Goal
Summarize paired x and y data with a correlation coefficient, least-squares line, fitted values, and a transparent prediction.
Worldwide context
Saved once here, used across the site.
Currency changes display only. Country selection guides tax input; no tax rate is guessed.
Summarize paired x and y data with a correlation coefficient, least-squares line, fitted values, and a transparent prediction.
Slope = Sxy/Sxx; intercept = ȳ − slope×x̄; r = Sxy/√(SxxSyy); fitted y = intercept + slope×x.A clearer path to an answer
This page keeps the calculation transparent: define the goal, enter the matching values, inspect the method, and decide what the result means in your situation.
Summarize paired x and y data with a correlation coefficient, least-squares line, fitted values, and a transparent prediction.
x values · Matching y values · x value to predict
Slope = Sxy/Sxx; intercept = ȳ − slope×x̄; r = Sxy/√(SxxSyy); fitted y = intercept + slope×x.
Calculate, review the assumptions below, then compare a related tool when the decision needs more context.
Summarize paired x and y data with a correlation coefficient, least-squares line, fitted values, and a transparent prediction.
Open the Scatter Plot and Linear Regression Calculator pageMore statistics tools
Download PDFDownload Word (.doc)
Enter your values above and choose Calculate to see the result here.
Calculation map
Slope = Sxy/Sxx; intercept = ȳ − slope×x̄; r = Sxy/√(SxxSyy); fitted y = intercept + slope×x.
Bounded, transparent calculation
Your recent runs stay in this browser session only.
Formula: Slope = Sxy/Sxx; intercept = ȳ − slope×x̄; r = Sxy/√(SxxSyy); fitted y = intercept + slope×x.
The calculator treats the two text lists as paired observations, computes the least-squares line, reports r and R², and returns fitted values for each row. The data table is meant to be read alongside the summary rather than treated as proof of causation.
Worked example: The fitted line is approximately y = 2x − 0.4 and predicts about 11.6 at x = 6.
The displayed limits are checked before the handler runs. Model-specific domain checks may also reject impossible or non-finite inputs.
Methodology: This calculator follows the WorldCalculate input, formula, precision, and boundary policy. Read the official methodology.
Calculator usage statistics
This section counts anonymous successful Calculate submissions, not unique visitors. Counts and top tools appear only when trusted aggregate data is available; country analysis is shown only under the same condition and reporting threshold.
Answer-first guide
Summarize paired x and y data with a correlation coefficient, least-squares line, fitted values, and a transparent prediction. Start with one clearly defined goal, enter values in the units shown, and keep the result attached to the assumptions below.
This tool is useful when your question includes scatter plot calculator, linear regression calculator, correlation coefficient. It returns the outputs declared in the calculator contract rather than a live quote, approval, diagnosis, or professional sign-off.
x values · Matching y values · x value to predict. Keep the same time period, unit system, and currency wherever the form requires comparable values.
Run the worked example first, compare its output with the page's example, then change one input at a time. This makes an unexpected result easier to trace to a unit, boundary, or assumption.
Need a wider view? Browse Statistics Calculators or compare the related tools below. The WorldCalculate methodology explains how formulas, examples, limits, and revisions are reviewed.
Slope = Sxy/Sxx; intercept = ȳ − slope×x̄; r = Sxy/√(SxxSyy); fitted y = intercept + slope×x.
The calculator treats the two text lists as paired observations, computes the least-squares line, reports r and R², and returns fitted values for each row. The data table is meant to be read alongside the summary rather than treated as proof of causation.
The fitted line is approximately y = 2x − 0.4 and predicts about 11.6 at x = 6.
Context and background
Statistics tools describe data or evaluate a stated probability model. They do not turn an observed summary into causation, certainty, or a forecast without additional evidence.
Data analysis developed from summaries of observations into probability, estimation, and decision measures. The essential habit remains the same: define the population, sample, variable, and convention before calculating.
Research and review
Researched by Hassan ALRowaie, Founder and editorial researcher at WorldCalculate.
This guide follows the live calculator's declared inputs, formula, worked example, assumptions, validation boundaries, and source-backed methodology. The review date describes editorial review of the calculator explanation; it is not a promise that external facts or rates remain current.
A scatter plot starts with paired observations: an x value and the y value measured with it. A regression line then summarizes the linear pattern, if one is appropriate. This calculator keeps those stages connected. It accepts two matching lists, computes the sample means and least-squares line, reports the correlation coefficient and R², and shows fitted values in a data table. The page is built for interpretation as much as arithmetic: a strong correlation is not proof that one variable causes the other, and a line can be a poor summary when the pattern is curved or dominated by an outlier.
The first x value belongs with the first y value, the second x with the second y, and so on. Changing the order of one list changes the relationships in the data. Use commas, spaces, or semicolons as separators, but do not include labels inside the lists.
At least three pairs are required for this worksheet. The x list must vary because a vertical collection of points has no ordinary slope for predicting y from x. The calculator checks lengths and finite values before doing the regression arithmetic.
The least-squares slope is Sxy/Sxx. It measures how much the fitted y value changes for one unit of x in the sample. The intercept is chosen so the line passes through the point of means (x̄, ȳ). Together they form y = slope×x + intercept.
The fitted value is not necessarily an observed data point. It is the value on the line at a particular x. The table shows observed and fitted y values side by side, making the residual—the observed value minus fitted value—conceptually visible even when the main output summarizes its squared total.
The correlation coefficient r describes the direction and strength of a linear association, with values near 1 or −1 indicating a tight positive or negative pattern and values near zero indicating little linear association. R² is r squared and summarizes the share of sample variation in y aligned with the linear fit under this model.
These numbers are summaries, not universal quality labels. A nonlinear relationship can have a modest r even when the variables are strongly related, and a small dataset can make the summary unstable. Always inspect the points and context.
The prediction input evaluates the fitted line at the selected x. A value inside the observed x range is interpolation; a value outside it is extrapolation. Extrapolation assumes the same pattern continues beyond the data, which may be unreasonable in physical, financial, or social settings.
The page does not attach a confidence interval, causal interpretation, or future guarantee to the prediction. If a decision matters, report the data range, uncertainty, sample design, and the reason a linear model was selected.
Start by plotting or otherwise inspecting the paired data. Then check the returned slope, r, R², and fitted table. Look for curvature, clusters, changing spread, and influential points. Finally ask whether the variables were measured in a way that supports the interpretation you want.
For a classroom exercise, calculate the line by hand for a small dataset and compare the slope and intercept. For a real analysis, preserve the original observations and document any cleaning or exclusions before relying on the summary.
A residual is the observed y value minus the fitted y value. Positive residuals sit above the line and negative residuals sit below it. A random-looking mix is often more compatible with a linear summary than a curved or fan-shaped residual pattern, although formal analysis requires more than this compact page.
One unusual point can change the slope and correlation noticeably, especially in a small dataset. Inspect the pairs instead of deleting a point simply because it makes the result inconvenient. Any exclusion should have a documented measurement or study reason.
The slope carries units: y units per x unit. Changing kilometres to metres changes the numerical slope even though the relationship is the same. The correlation coefficient is unit-free, but it still depends on the paired observations and their linear pattern.
Call x the explanatory variable only when that role is justified by the question. Swapping x and y changes the regression prediction direction and slope interpretation, even though the correlation coefficient itself has the same sign and magnitude.
Does a high r prove causation? No. Two variables can move together because of a third factor, shared time trend, selection effect, or chance. Does R² mean the model is correct? It summarizes sample fit under a linear model, but it cannot prove the model is appropriate or future-stable.
Why can the prediction be unreliable? A line is an approximation and uncertainty grows when the input is far from the observed range. The page reports the arithmetic result while leaving confidence intervals and domain-specific validation to a fuller analysis.
A regression can be numerically correct and still answer the wrong question. Decide whether one row represents a person, day, machine cycle, school, product, or another unit. Then identify which variable is the candidate input and which is the outcome. Mixing rows from different units can create an association that does not exist at the level you care about.
Write the question in one sentence before entering data: for example, ‘How does measured travel time change with distance on these trips?’ That sentence determines the roles of x and y, the units, the relevant population, and whether a prediction is meaningful. A clear question makes the output easier to explain and harder to overstate.
Each pair should come from the same case and measurement frame. If x is a student’s study hours and y is that student’s score, do not sort the lists independently or pair average hours from one class with scores from another. The pairing is the data story that the fitted line summarizes.
Keep a row identifier outside the calculator when the data matter. Record when and how each value was measured, who was included, and whether any observations were excluded. The text-list interface is convenient for a small example, but it is not a substitute for a source dataset or data dictionary.
A scatter plot can reveal a curve, a cluster, a gap, a ceiling, or a single influential point before any formula is applied. If the points form a U-shape, a line may report a weak correlation even though x and y are strongly related. If two groups have different colors or origins, one combined line may hide the structure.
Use this calculator after a visual inspection, not as the first and only analysis. The returned fitted values are easier to interpret when you know what the points look like. A nearby chart, residual plot, or group summary can provide information that one slope and one correlation cannot.
The slope is the fitted change in y for a one-unit increase in x. If x is minutes and y is kilometres, the slope has units of kilometres per minute. Changing x from minutes to hours changes the numerical slope but not the fitted line when the conversion is done consistently.
A slope is local to the observed data and model. It does not say that every individual case changes by exactly that amount, and it does not establish a mechanism. Report the unit, observed x range, and whether the change is a descriptive average or a causal effect.
The intercept is the fitted y value when x equals zero. It is required to position the least-squares line, but x = 0 may be outside the observed data or impossible in the real setting. A negative intercept can be mathematically valid even when negative y values are not physically possible in the application.
Do not interpret the intercept as a baseline cause without checking the question. Centering x can make the intercept represent the fitted outcome at a typical x, while keeping the original slope interpretation through a documented transformation. The calculator reports the ordinary line; context determines which coefficient deserves a narrative.
R² is the square of the correlation reported by this calculator. In a simple linear fit with an intercept, it summarizes how closely the observed y values align with the fitted line in the entered sample. It is conditional on the variables, observations, scale, and model form supplied.
A high R² can occur in a biased sample, a time trend, or a relationship that will not continue outside the study. A lower R² does not make a model useless if the outcome is noisy or the goal is descriptive. Pair R² with a plot, residual review, sample design, and a statement of purpose.
The difference between observed and fitted y is a residual. Listing fitted values lets you compute and inspect those differences. Residuals that grow with x can suggest changing spread; a curved sequence can suggest that a line misses systematic structure; a block of large residuals can signal a group or measurement issue.
A residual pattern does not automatically select a better model. It tells you where to ask the next question: transform a variable, add a justified predictor, separate groups, collect more data, or retain the line as a limited descriptive summary. Do not conceal a visible pattern by reporting only R².
An outlier has an unusual y value relative to the pattern; a high-leverage point has an unusual x value. A point can be both. Such observations may be errors, rare but valid cases, or the most informative part of the question. The calculator cannot determine which one it is.
Run a documented sensitivity check with and without a questionable point only when the data and study rules permit it. Report the reason for any exclusion and show how the line changes. Deleting a point because it weakens a desired story is not data cleaning; it is a change to the evidence.
Several subgroups can create a strong overall correlation even when the within-group relationships are weak or reversed. This can happen when group membership is related to both x and y. Add group labels to the source data and inspect the groups before treating the combined line as a general law.
Averaging observations can also change the relationship. A line through group means answers a different question from a line through individual rows. Keep the level of aggregation visible and do not use an aggregate prediction as though it described every member of the underlying population.
Interpolation evaluates the line inside the observed x range, where the sample provides some support for the input region. Extrapolation evaluates it outside that range and assumes the relationship continues. Physical limits, policy changes, saturation, and new populations can make that assumption fail quickly.
Write the minimum and maximum observed x values beside any prediction. If the chosen x is outside them, label the answer extrapolated and identify the additional evidence needed. A prediction at x = 6 from data spanning 1 to 5 is not equivalent to a measured observation at 6.
The calculator returns a fitted point prediction, not a confidence interval for the mean response or a prediction interval for a new observation. Those intervals require residual variation, sample size, model assumptions, and a stated confidence level. The point on the line should therefore not be presented as an exact future value.
For a decision, add uncertainty from the appropriate statistical method and explain what the interval covers. A narrow interval around the mean does not mean an individual future result will be close to the line. Distinguish uncertainty in the fitted relationship from uncertainty in the measurement or future context.
If the purpose is prediction, separate data used to fit the line from data used to check it when the dataset is large enough. A line that fits the same observations can look better than it performs on new cases. For a small classroom example, say explicitly that no independent validation set is available.
Communicate the result in a compact order: describe the observations, state the fitted equation with units, give r and R² as conditional summaries, show the data range, report the chosen prediction, and name the main limitation. This sequence answers both ‘what did the calculator do?’ and ‘why should I care?’
A correlation coefficient summarizes paired numbers; it does not tell you how the numbers were assigned or whether alternative explanations were controlled. An observational association can reflect confounding, reverse direction, selection, common time movement, or measurement bias. A randomized experiment and a regression line answer different questions.
Use causal language only when the study design and domain evidence support it. Safer wording is often ‘in this sample, y tended to increase as x increased’ followed by the range and limitations. Careful language does not weaken a result; it tells readers exactly what the data support.
Save the original x and y lists, the calculator inputs, the returned equation, r, R², fitted values, and the date of analysis. Add the source, unit definitions, cleaning decisions, prediction value, and observed range. Someone reviewing the work should be able to recreate the result without guessing how the rows were paired.
Finish with a decision statement: use the line for a limited description, collect more observations, investigate an outlier, fit a different model, or obtain domain review. The calculator is most useful when it leads to a better next question instead of ending the analysis at a single coefficient.
A curved or strongly uneven relationship may be easier to describe after a justified transformation, such as a logarithm, square root, or reciprocal. A transformation changes the meaning of the slope and prediction, so it must be recorded rather than treated as a cosmetic change. This calculator fits the values exactly as entered and does not choose a transformation automatically.
If x is transformed, say whether the prediction is on the transformed scale or has been converted back to the original units. Do not transform only one list because it makes r look stronger without explaining the scientific or practical reason. A better fit is useful only when the transformed question remains meaningful to the reader.
Blank cells, text labels, sentinel values, and impossible measurements should be resolved before creating the two lists. Removing one value from only one list breaks the pairing; removing a row from both lists changes the sample and must be documented. The calculator checks finite numeric input but cannot know whether a value is a legitimate measurement.
Keep a cleaning log with the original row count, excluded rows, reason, and person who approved the decision. If missingness is related to x or y, the remaining pairs may not represent the original population. A tidy list can therefore be less trustworthy than a slightly messy dataset whose limitations are openly described.
Three pairs are enough for the calculator to perform its arithmetic, but they are rarely enough to support a stable general conclusion. With a small sample, one observation can dominate the slope and r, and the fitted line may change substantially when another case is added. Treat small examples as demonstrations or preliminary evidence.
Generalization also depends on how cases were selected. A volunteer sample, a single classroom, one season, or one machine may not represent the wider population. State the population that the data actually cover, and avoid describing a sample-specific line as a universal rule without external validation.
Before publishing a regression result, confirm the paired rows, variable roles, units, sample source, observed x range, fitted equation, r, R², residual pattern, outlier treatment, and prediction status. Mark interpolation and extrapolation differently. If a transformation, aggregation, or exclusion was used, explain it near the result rather than burying it in a footnote.
End with what the reader should do next: use the calculator to reproduce the example, collect a better sample, inspect a plot, validate on new data, or ask a subject expert. A strong scatter-plot article leaves the reader able to challenge the line as well as use it. Record that next action with the result so the analysis has a clear owner and follow-up. A short provenance note also prevents a later reader from confusing a classroom illustration with a validated production forecast. This remains reproducible.
Summarize paired x and y data with a correlation coefficient, least-squares line, fitted values, and a transparent prediction.
Slope = Sxy/Sxx; intercept = ȳ − slope×x̄; r = Sxy/√(SxxSyy); fitted y = intercept + slope×x. The calculator treats the two text lists as paired observations, computes the least-squares line, reports r and R², and returns fitted values for each row. The data table is meant to be read alongside the summary rather than treated as proof of causation.
Enter x values, Matching y values, x value to predict, then choose Calculate.
x and y lists contain the same number of paired finite values. At least three observations are supplied. x values contain variation so Sxx is positive. The fitted line summarizes linear association in the entered sample. A prediction outside the observed x range is extrapolation and needs extra justification. Correlation does not establish causation, and outliers can strongly affect the line.
This calculator is part of the WorldCalculate library. Its formula, example, assumptions, input bounds, and output formatting follow the official methodology.
These WorldCalculate collections connect this tool with related questions while keeping each calculation separate and transparent.