Back to the on-screen lesson ·

Fitted values, residuals and R-squared

Split total variation into explained and residual sums of squares, compute R-squared, and say what it does not measure.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute fitted values, residuals and the sums of squares, find R-squared two ways, and judge what a given R-squared does and does not show.

2. What you already have

From the last lesson you can fit an OLS line by hand and you know its residuals sum to zero and are uncorrelated with $x$. From statistics you know the variance as an average squared deviation from the mean.

This lesson asks how well the line fits. It splits the total variation in $y$ into the part the line explains and the part it leaves in the residuals, and turns the split into one number, $R^2$ — then explains why that number answers much less than people expect.

3. Terms to use precisely

TermWhat it means
Fitted value$\hat y_i = \hat\beta_0 + \hat\beta_1 x_i$, the point on the line above $x_i$.
Residual$\hat u_i = y_i - \hat y_i$, the vertical gap from the point to the line.
SSTTotal sum of squares, $\sum (y_i - \bar y)^2$: all the variation in $y$.
SSEExplained sum of squares, $\sum (\hat y_i - \bar y)^2$: variation in the fitted values.
SSRSum of squared residuals, $\sum \hat u_i^2$: variation the line leaves over.
R-squared$R^2 = SSE/SST = 1 - SSR/SST$, the share of sample variation in $y$ explained.
Goodness of fitHow closely the data points lie to the fitted line.

4. Total variation = explained + residual

Every observation can be split into a fitted part and a residual: $y_i = \hat y_i + \hat u_i$. Subtract $\bar y$ from both sides and square. Because OLS residuals are uncorrelated with the fitted values, the cross term sums to zero, and the squares add up cleanly:

$$\underbrace{\sum (y_i - \bar y)^2}_{SST} = \underbrace{\sum (\hat y_i - \bar y)^2}_{SSE} + \underbrace{\sum \hat u_i^2}_{SSR}.$$

The share explained is $R^2 = SSE/SST = 1 - SSR/SST$. It is $0$ when the line is flat and explains nothing, $1$ when every point sits on the line, and in between otherwise. With one regressor it equals the squared correlation of $x$ and $y$.

Textbooks disagree about the letters: some call the explained sum SSR, for regression, and the residual sum SSE, for error. The meaning is what matters: always check which one is left over.

Another way: action

On a scatterplot, draw the horizontal line at $\bar y$ and the fitted line. For one point, mark the vertical distance to the mean (total), from the mean to the line (explained) and from the line to the point (residual).

Another way: steps

  1. Fit the line and compute each fitted value.
  2. Subtract to get each residual, and check they sum to zero.
  3. Square and sum the residuals for SSR.
  4. Square and sum the deviations of $y$ from its mean for SST.
  5. Divide: $R^2 = 1 - SSR/SST$.

5. Why the cross term vanishes

Write $y_i - \bar y = (\hat y_i - \bar y) + \hat u_i$ and square it. You get the two squares plus twice the cross product $\sum (\hat y_i - \bar y)\hat u_i$.

That cross product is zero because of the normal equations from the last lesson. The residuals sum to zero, which kills the $\bar y$ part, and the residuals are uncorrelated with $x$, so they are uncorrelated with any linear function of $x$, including the fitted values. So OLS splits the variation into two pieces that do not overlap.

This is special to least squares with an intercept. Fit a line some other way, or force it through the origin, and the pieces no longer add up; $R^2$ computed the two ways can even disagree.

6. What R-squared does not tell you

$R^2$ answers one question: how closely do these sample points cluster around this line? It is silent on the questions a policy analyst usually cares about.

In cross-section data on people, $R^2$ of $0.1$ to $0.3$ is normal and says only that individuals differ in many ways the regression does not measure. In time series of trending variables, $R^2$ near $1$ is common and often meaningless.

7. Units and scale

Residuals and the three sums of squares are in the units of $y$ (squared, for the sums). $R^2$ has no units: it is a ratio of two sums of squares measured the same way.

That makes it convenient for comparing fits of the same $y$ with different regressors, but dangerous for comparing across different outcome variables. A regression of log wages and a regression of wages can have very different $R^2$ with the same information, because the total variation being explained is different. Compare $R^2$ only between models of the same dependent variable, estimated on the same sample.

Residuals, by contrast, keep the units of the outcome and are worth reading one by one. A residual of $-3$ in a sales regression says that week sold three thousand units fewer than weeks with the same advertising did on average. Sorting observations by their residuals is a quick way to find data errors, unusual cases and patterns the line has missed, such as every holiday week sitting well above it. A pattern in the residuals is a hint that the model needs another variable or a different shape, and it is often the first clue an analyst gets.

8. Adding regressors always raises R-squared

In the next unit, regressions get more than one explanatory variable. One warning belongs here, because it is about fit. Adding any variable to a regression, even a column of random numbers, can never lower $R^2$ and almost always raises it a little. The reason is mechanical: with an extra coefficient to choose, least squares can do at least as well as before by setting it to zero, and usually finds some small value that shaves a bit off the residuals by chance.

So a higher $R^2$ from a bigger model is not evidence that the new variable matters. The adjusted $R^2$ charges a penalty for each extra coefficient, and can fall when a useless variable is added. Better still, later lessons test whether new variables matter with an F-test, which asks whether the improvement in fit is larger than chance alone would produce.

The deeper lesson is that fit is cheap. With enough variables, any dataset can be fitted closely. What is expensive, and what an econometrician is paid for, is a credible argument that a particular coefficient answers a particular question.

9. Prediction versus explanation

When the goal is prediction, fit is exactly the right target, but it must be measured on data the model has not seen. A model that fits its own sample with $R^2 = 0.95$ can predict new cases badly if it has chased noise. Forecasters therefore hold back part of the data, fit on the rest, and judge the model by how well it predicts the held-back part.

When the goal is a causal effect, fit is almost beside the point. The question is whether the variation in $x$ used to estimate the slope is free of selection. A randomized experiment with $R^2 = 0.02$ can give a far more trustworthy effect than an observational regression with $R^2 = 0.8$.

10. Working a fit, step by step

Five points, y against x: (1, 6), (2, 5), (3, 9), (4, 13) and (5, 12), with the least squares line ŷ = 3 + 2x. The fitted values are 5, 7, 9, 11 and 13, so the residuals are 1, −2, 0, 2 and −1: two points above the line, two below and one on it. The residuals sum to zero, as they always do with an intercept.
Five points, y against x: (1, 6), (2, 5), (3, 9), (4, 13) and (5, 12), with the least squares line ŷ = 3 + 2x. The fitted values are 5, 7, 9, 11 and 13, so the residuals are 1, −2, 0, 2 and −1: two points above the line, two below and one on it. The residuals sum to zero, as they always do with an intercept.

The figure plots these five points and the fitted line; each residual is the vertical gap between a point and the line.

Take five points $x = 1, 2, 3, 4, 5$ and $y = 6, 5, 9, 13, 12$ with OLS line $\hat y = 3 + 2x$.

  1. Fitted values. $5, 7, 9, 11, 13$.
  2. Residuals. $1, -2, 0, 2, -1$, which sum to zero as they must.
  3. SSR. $1 + 4 + 0 + 4 + 1 = 10$.
  4. SST. $\bar y = 9$, deviations $-3, -4, 0, 4, 3$, squares summing to $50$.
  5. R-squared. $1 - 10/50 = 0.8$.

The line explains eighty percent of the variation in $y$ in these five points. As a check, the fitted values' squared deviations from $9$ are $16 + 4 + 0 + 4 + 16 = 40$, and $40/50 = 0.8$ too.

11. How to check a fit answer

Run three checks on every calculation.

  1. Do the residuals sum to zero? With an intercept they always do. If not, a fitted value is wrong.
  2. Does SSE plus SSR equal SST? Compute all three and add. A mismatch means a slip in one of the sums.
  3. Is $R^2$ between zero and one? A value outside that range means SSR and SST were swapped or a sum is wrong.

Then an interpretation check: have you described $R^2$ as a share of sample variation, and not as evidence that the slope is causal or correctly estimated?

12. In the world: how much risk is the market's?

Financial analysts regress a stock's monthly return on the return of the whole market, such as the S&P 500. The slope is the stock's beta, and the $R^2$ has a direct meaning: the share of the stock's return variation that comes from market-wide movements rather than from news about the company itself.

Suppose a utility company's returns have a total sum of squares of $400$ and the regression leaves $160$ in the residuals. Then $R^2 = 0.6$: sixty percent of the stock's variation moves with the market, and forty percent is company-specific. For a small biotech firm, $R^2$ might be only $0.1$, because its price is driven mainly by trial results and approvals.

That forty or ninety percent of company-specific variation is the risk an investor can remove by holding many stocks, since company news tends to cancel out across a portfolio. Market risk cannot be diversified away. Here $R^2$ is being used for exactly what it measures, a split of variation, and not as a claim about cause.

Portfolio managers use the same arithmetic in reverse. A fund that claims to pick stocks cleverly but whose returns have an $R^2$ of $0.98$ against the index is, in practice, charging a fee to hold the index.

13. A high R-squared is not a good causal estimate

The most common mistake is reading $R^2$ as a measure of how right the model is. It measures how closely the points fit a line; a model can fit closely and be causally wrong, or fit loosely and be causally right.

A second mistake is thinking a low $R^2$ makes a slope useless. With enough data, a slope can be estimated precisely even when most of the variation in $y$ comes from other sources.

A third is computing the residual as fitted minus actual. The convention is actual minus fitted, so points above the line have positive residuals.

14. R-squared from two sums

  1. Read the total sum of squares.

    $SST = 200$

    All the variation in $y$.

  2. Read the residual sum.

    $SSR = 50$

    What the line misses.

  3. Form the residual share.

    $50/200 = 0.25$

    A quarter unexplained.

  4. Subtract from one.

    $R^2 = 0.75$

    Three quarters explained.

  5. Check with the explained sum.

    $SSE = 150, \ 150/200 = 0.75$

    Both routes agree.

15. Residuals for a three-point fit

  1. Write the line.

    $\hat y = 4 + 3x$

    Given by OLS.

  2. Compute the fitted values.

    $7, 10, 13$

    At $x = 1, 2, 3$.

  3. Read the actual values.

    $6, 12, 12$

    The data.

  4. Subtract fitted from actual.

    $-1, 2, -1$

    The residuals.

  5. Check their sum.

    $-1 + 2 - 1 = 0$

    Always zero with an intercept.

  6. Sum their squares.

    $1 + 4 + 1 = 6$

    The SSR.

16. A full decomposition

  1. Write the data and line.

    $x: 0..4; \ y: 7,15,16,15,17; \ \hat y = 10 + 2x$

    Five points.

  2. Compute the fitted values.

    $10, 12, 14, 16, 18$

    On the line.

  3. Compute the residuals.

    $-3, 3, 2, -1, -1$

    They sum to zero.

  4. Sum the squared residuals.

    $9 + 9 + 4 + 1 + 1 = 24$

    SSR.

  5. Compute SST about the mean.

    $\bar y = 14; \ 49 + 1 + 4 + 1 + 9 = 64$

    Total variation.

  6. Compute the R-squared value.

    $1 - 24/64 = 0.625$

    The explained share.

  7. Check with SSE.

    $16 + 4 + 0 + 4 + 16 = 40 = 64 - 24$

    The pieces add up.

17. Your turn: SST = 400 and SSR = 160.

  1. Subtract the residual sum.

    $SSE = 400 - 160 = 240$

    Explained variation.

  2. Divide by the total.

    $240/400 = 0.6$

    The explained share.

  3. Your turn: work this step out. Its working is at the end of the packet.

    State the result.

18. Guided practice

An OLS line is $\hat y = 15 + 2x$. One observation has $x = 2$ and $y = 22$. What is its residual?

Answer: residual

19. Guided practice

Complete the worked solution: a regression has $SST = 460$ and $SSR = 322$. Find the explained sum of squares, the residual share and the $R^2$.

  1. Subtract the residual sum.

    $SSE = 460 - 322 =$ a

    What the line explains.

  2. Form the residual share.

    $SSR / SST =$ b

    The fraction left unexplained.

  3. Take one minus that share.

    $R^2 =$ c

    The explained fraction.

20. Guided practice

A regression has total sum of squares $SST = 120$ and sum of squared residuals $SSR = 90$. What is its $R^2$?

Answer: R-squared

21. Practice

In a simple regression, the correlation between $x$ and $y$ is $0.9$. What is the regression's $R^2$?

Answer: R-squared

22. Practice

The OLS line for three points is $\hat y = 7 + 5x$. The points are $(1, 11)$, $(2, 19)$ and $(3, 21)$. Fill in each fitted value and residual.

fitted valueresidual
x = 1
x = 2
x = 3

23. Practice

Match each claim about a regression with $R^2 = 0.9$ to a verdict.

follows from R² = 0.9does not follow from R² alone
The line accounts for 90% of the sample variation in y.
x causes 90% of the changes in y.
Residuals are small relative to the spread of y.
The slope is free of omitted-variable bias.

24. Somewhere new

An analyst regresses a company's monthly stock return on the market's return. The total sum of squares of the company's returns is $500$ and the sum of squared residuals is $300$. What share of the company's return variation does the market explain?

Answer: share of variation

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

The OLS line for three points is $\hat y = 5 + 2x$. The points are $(1, 4)$, $(2, 15)$ and $(3, 8)$. Fill in each fitted value and residual.

fitted valueresidual
x = 1
x = 2
x = 3

27. What you can do now

You can measure a regression's fit. Explain to someone why a regression with an R-squared of 0.9 can still give a misleading answer to a causal question.

Working for the steps left to you

17. Your turn: SST = 400 and SSR = 160., step 3

$R^2 = 0.6$

Sixty percent of the variation.