Linear Regression Calculator
Find the least-squares regression line from X and Y data: slope, intercept, R², r, predictions, residuals, a scatter plot and every calculation step.
- Slope (b₁)
- 2
- Y-intercept (b₀)
- −0.2
- R² (coefficient of determination)
- 0.9804
- 98.0392% of the variation in Y is accounted for by the line, in this dataset
- Correlation (r)
- 0.990
- n = 5 pairs
If X increases by 1 unit, the model predicts that Y will increase by about 2 units on average. This describes the fitted line, not a proven cause and effect.
The y-intercept −0.2 is the model's predicted Y when X = 0. X = 0 is outside your observed X range (1 to 5), so the intercept may have no meaningful real-world interpretation; it mainly positions the line.
Predicted Y = 6.8
ŷ = b₀ + b₁x
ŷ = −0.2 + 2 × 3.5
ŷ = 6.8
X = 3.5 is within your observed range (1 to 5): an interpolation.
Scatter plot and regression line
ŷ = −0.2 + 2x
R² = 0.9804
Dots are your observed points; the dashed line is the fitted least-squares line; the diamond is your prediction.
Dotted vertical lines are the residuals, e = y − ŷ: the distance from each point to the line. Least squares chooses the line that makes the sum of their squares as small as possible.
The fitted line has a positive slope, so the model predicts higher Y values as X increases. With R² = 0.980, the observed points lie fairly close to the line in this dataset. Whether that is useful depends on the context and purpose.
Residual plot
- Random scatter around zero is consistent with a straight-line model.
- A curved pattern (for example a U or arch) suggests the relationship is not linear.
- A funnel shape, with spread growing or shrinking, suggests non-constant variance.
- One point far from the rest may have a large influence on the line.
- These are visual checks, not formal diagnostic tests.
Predicted values and residuals
Residual e = y − ŷ: positive when the point is above the line, negative when below. For a least-squares line with an intercept, the residuals add up to 0.
| # | X | Observed Y | Predicted ŷ | Residual e |
|---|---|---|---|---|
| 1 | 1 | 2 | 1.8 | 0.2 |
| 2 | 2 | 4 | 3.8 | 0.2 |
| 3 | 3 | 5 | 5.8 | −0.8 |
| 4 | 4 | 8 | 7.8 | 0.2 |
| 5 | 5 | 10 | 9.8 | 0.2 |
How the regression line was calculated
b₁ = Σ[(x − x̄)(y − ȳ)] / Σ[(x − x̄)²]
b₀ = ȳ − b₁x̄
ŷ = b₀ + b₁x
| # | X | Y | x − x̄ | y − ȳ | (x − x̄)(y − ȳ) | (x − x̄)² |
|---|---|---|---|---|---|---|
| 1 | 1 | 2 | −2 | −3.8 | 7.6 | 4 |
| 2 | 2 | 4 | −1 | −1.8 | 1.8 | 1 |
| 3 | 3 | 5 | 0 | −0.8 | 0 | 0 |
| 4 | 4 | 8 | 1 | 2.2 | 2.2 | 1 |
| 5 | 5 | 10 | 2 | 4.2 | 8.4 | 4 |
| Σ | 15 | 29 | 0 | 0 | 20 | 10 |
Step 1: Means
n = 5
x̄ = Σx / n = 15 / 5 = 3
ȳ = Σy / n = 29 / 5 = 5.8
Step 2: Sums of deviations (see the table)
Σ(x − x̄)(y − ȳ) = 20
Σ(x − x̄)² = 10
Step 3: Slope
b₁ = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)²
b₁ = 20 / 10 = 2
Step 4: Intercept
b₀ = ȳ − b₁x̄
b₀ = 5.8 − 2 × 3 = −0.2
Step 5: Regression equation
ŷ = −0.2 + 2x
Step 6: R² from the residuals
SSE is the sum of squared residuals (y − ŷ)²; SST is the total variation of Y around its mean.
SSE = 0.8, SST = Σ(y − ȳ)² = 40.8
R² = 1 − SSE / SST = 1 − 0.8 / 40.8 = 0.9804
Regression equation
- ŷ
- predicted value of Y
- b₀
- y-intercept: predicted Y when x = 0
- b₁
- slope: predicted change in Y per 1-unit increase in x
- x
- input (predictor) value
- e
- residual, y − ŷ
In simple least-squares regression with one predictor and an intercept, R² equals r², the square of the Pearson correlation. That is not true of every kind of regression.
| Correlation | Linear regression |
|---|---|
| Measures strength and direction of a linear association | Fits a predictive straight-line equation |
| Produces r | Produces a slope and intercept (and R²) |
| No prediction equation | Estimates Y from X |
| Symmetric: swapping X and Y gives the same r | X is the predictor, Y the response; swapping them changes the equation |
Calculate Pearson correlation separately, with an interactive demo, using the Correlation Calculator. Summarize each variable with the Average Calculator and the Standard Deviation Calculator; estimate a mean with a margin of error in the Confidence Interval Calculator; and plan how much data to collect with the Sample Size Calculator.
A linear regression calculator finds the straight line that best fits paired data and turns it into an equation you can use for predictions. For X = 1, 2, 3, 4, 5 and Y = 2, 4, 5, 8, 10, the least-squares line is ŷ = −0.2 + 2x, with R² = 0.980.
Enter your X and Y values to get the regression equation, the slope and y-intercept with plain-language interpretations, R² and the correlation coefficient r, and a prediction for any X you choose. The calculator plots the data with the line of best fit and the residuals, draws a residual plot, and shows the full calculation table and formulas so you can check the work.
Worked Calculation Examples
| Scenario | Result | Calculation Step |
|---|---|---|
| Study hours vs exam score (hypothetical) | ŷ = 51.087 + 2.860x, R² = 0.902 | Each extra hour of study is associated with about 2.86 more points on average. For 7 hours the model predicts ŷ = 51.087 + 2.860 × 7 ≈ 71.1 (an interpolation, since 7 is between 2 and 12). R² ≈ 0.902: the line accounts for about 90% of the variation in scores in these 8 students. It does not show that studying caused the scores. |
| Advertising spend vs sales (hypothetical, $ thousands) | ŷ = 11.071 + 3.012x, R² = 0.987 | The model predicts about $3,012 more sales for each extra $1,000 of advertising, and about $11,071 of sales with no advertising (X = 0 is just outside the data, so treat that with care). At $5,500 of spend, ŷ ≈ 11.071 + 3.012 × 5.5 ≈ 27.64, or about $27,640. Predicting for $50,000 of spend would be an extrapolation far beyond the data. |
| A weak linear relationship: car age vs owner rating (hypothetical) | ŷ = 6.267 + 0.006x, R² ≈ 0.0002 | The fitted line is almost flat and R² is essentially 0: the linear model explains almost none of the variation in ratings, so its predictions are barely better than using the average rating of 6.3 for every car. This does not rule out a non-linear relationship or other factors. |
What Is Linear Regression?
Linear regression is a statistical method that models the relationship between a response variable Y and a predictor variable X with a straight line. Simple linear regression, with one predictor, has the form ŷ = b₀ + b₁x, where ŷ is the predicted value of Y, b₀ is the y-intercept, b₁ is the slope and x is the input value.
It is used to summarize how Y tends to change with X, and to estimate Y for values of X, for example predicting sales from advertising spend or exam scores from study hours.
How Linear Regression Works
For any line, each observed point has a residual: the vertical distance e = y − ŷ between the point and the line. The least-squares regression line is the one that makes the sum of squared residuals, Σe², as small as possible, which is why it is called the line of best fit.
1. Collect paired X and Y data.
2. Calculate the means x̄ and ȳ.
3. Calculate the slope b₁ = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)².
4. Calculate the intercept b₀ = ȳ − b₁x̄.
5. Write the equation ŷ = b₀ + b₁x.
6. Calculate the predicted values and residuals.
7. Examine the residuals and R² to judge the fit.
For X = 1 to 5 and Y = 2, 4, 5, 8, 10: x̄ = 3, ȳ = 5.8, Σ(x − x̄)(y − ȳ) = 20 and Σ(x − x̄)² = 10, so b₁ = 20 / 10 = 2, b₀ = 5.8 − 2 × 3 = −0.2, and ŷ = −0.2 + 2x.
How to Interpret the Slope and Intercept
The slope b₁ is the predicted change in Y for a 1-unit increase in X. A slope of 2 means that, on average, the model predicts Y to be 2 units higher for each extra unit of X; a negative slope means Y is predicted to fall as X rises. This describes the fitted line, not a proven effect of X on Y.
The y-intercept b₀ is the model's predicted Y when X = 0. It only has a real-world meaning if X = 0 is plausible and inside, or near, the observed data. If your X values run from 20 to 60, the intercept is an extrapolation that mainly positions the line.
How to Interpret R²
R², the coefficient of determination, is the proportion of the variation in the observed Y values that is accounted for by the fitted line, within this dataset. It runs from 0 to 1: R² = 0.84 means 84% of the variation in Y around its mean is accounted for by the linear model, and 16% remains in the residuals.
R² is not a measure of cause and effect, and there is no single cut-off for a good R². In a physics experiment 0.9 may be poor, while in social science 0.3 can be useful. A high R² also does not prove a straight line is the right model; a curved pattern can have a high R² and still be fitted badly, which the residual plot reveals.
For simple least-squares regression with an intercept, R² = r², the square of the Pearson correlation coefficient.
Correlation vs Regression
Correlation measures the strength and direction of a linear association with a single number, r, and treats X and Y symmetrically: swapping them gives the same r. Regression fits an equation that predicts Y from X. It treats X as the predictor and Y as the response, so regressing X on Y gives a different line. Use correlation to describe how closely two variables move together, and regression when you want to estimate one from the other.
Regression Does Not Prove Causation
A regression line can describe an association and make useful predictions within the range of the data, but fitting a line does not by itself show that changes in X cause changes in Y. Watch out for confounding variables that drive both, reverse causation, selection effects in how the data were collected, extrapolation beyond the data, non-linear relationships forced into a straight line, and outliers that pull the line.
Interpolation vs Extrapolation
Interpolation is predicting Y for an X value inside the range of your observed X values; extrapolation is predicting outside it. Extrapolation relies on the straight-line pattern continuing where you have no data, which it often does not: a plant's growth over the first 10 weeks does not continue in a straight line for years. The calculator still gives the prediction but warns you when the X value is outside the observed range.
Residuals, Outliers and Model Fit
The residual e = y − ŷ is the difference between an observed Y value and the value predicted by the regression line. Plotting residuals against X is a quick check of the model. Random scatter around zero is consistent with a straight line; a curved pattern suggests the relationship is not linear; a funnel shape suggests the spread changes with X; and a single point far from the rest may be pulling the line.
Outliers can substantially change the slope, the intercept and R². The calculator flags any point whose residual is more than 2 standard errors of the estimate from the line (for 6 or more points) and shows what the slope would be without it. Check such points for errors rather than deleting them automatically.
How to Use the Linear Regression Calculator
- Enter the X values (predictor) and Y values (response), separated by commas, spaces or new lines. The first X pairs with the first Y, and so on.
- Read the regression equation ŷ = b₀ + b₁x, the slope b₁, the y-intercept b₀, R² and the correlation coefficient r, with plain-language interpretations.
- Enter an X value under "Predict Y" to get ŷ with the substitution shown; the calculator warns you when the X value is outside your data (extrapolation).
- Check the scatter plot, where dotted lines show each residual (observed minus predicted), and the residual plot for curves or changing spread.
- Verify the result with the prediction table and the calculation table of deviations from the means, and follow the step-by-step formulas.
Frequently Asked Questions
What is a linear regression calculator?
A tool that fits the least-squares straight line to paired X and Y data and returns the regression equation ŷ = b₀ + b₁x, with the slope, intercept, R² and predictions.
How do you calculate linear regression?
Find the means x̄ and ȳ, compute the slope b₁ = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)², then the intercept b₀ = ȳ − b₁x̄. The regression equation is ŷ = b₀ + b₁x.
What is the regression equation?
ŷ = b₀ + b₁x: the predicted Y equals the intercept plus the slope times X. For example, ŷ = −0.2 + 2x.
What do slope and intercept mean?
The slope is the predicted change in Y for each 1-unit increase in X. The intercept is the predicted Y when X = 0, which is only meaningful if X = 0 is within or near the observed data.
What is R² in linear regression?
The proportion of the variation in the observed Y values accounted for by the fitted line, within the dataset, from 0 to 1. It is not a measure of causation. In simple linear regression, R² = r².
What is a residual?
The difference between an observed Y value and the value predicted by the line: e = y − ŷ. Positive residuals are above the line, negative ones below.
How do you predict Y from X?
Substitute X into the regression equation. With ŷ = −0.2 + 2x and X = 3.5: ŷ = −0.2 + 2 × 3.5 = 6.8.
What is the difference between correlation and regression?
Correlation gives one number, r, for the strength and direction of a linear association and is symmetric in X and Y. Regression fits an equation to predict Y from X, and swapping X and Y changes the equation.
Does linear regression prove causation?
No. It describes an association and can predict, but confounding variables, reverse causation or selection effects can produce the same pattern. Causal claims need study design, such as randomized experiments.
What is the difference between interpolation and extrapolation?
Interpolation predicts Y for an X inside the observed range; extrapolation predicts outside it and assumes the straight-line pattern continues where there is no data, so it is less reliable.
How many data points are needed for linear regression?
Two points define a line, but then the fit is always perfect and tells you nothing. In practice you want many more; with fewer than about 10 points the slope and R² can change a lot if one value changes.
Related Calculators
Correlation Calculator
Calculate the Pearson correlation coefficient (r) from paired X and Y data, with its strength, direction, r², a scatter plot and the full calculation.
Standard Deviation Calculator
Calculate sample and population standard deviation and variance of a dataset, with the mean, a deviation table and step-by-step working.
Average Calculator
Calculate the average (mean) of a list of numbers with the median, mode and range, or a weighted average, with step-by-step working.
Confidence Interval Calculator
Calculate a confidence interval for a population mean from the sample mean, standard deviation and sample size, using t or z, with the margin of error and every step.
Sample Size Calculator
Find the minimum sample size for a survey from the confidence level, margin of error, expected proportion and optional population size, with every step shown.
Z-Score Calculator
Calculate a z-score from a value, mean and standard deviation, with the steps, what it means, a bell-curve chart and the normal-distribution percentile.
Last updated: September 27, 2026.