Back to the on-screen lesson ·
The exact split of the response's variation into the part a line accounts for and the part it leaves, the fraction that results, and the four things that fraction is not.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will split a total sum of squares into explained and residual parts, compute the fraction of the variation a fitted line accounts for, and rank fits that share a total. You will also state precisely what that fraction does not establish about shape, about cause, or about the points themselves.
The last lesson split each observation into a fitted part and a residual. This one adds the squares of those pieces across the whole data set and asks what share of the response's variation the line has accounted for. The identity it rests on is the same decomposition, squared and summed.
The total sum of squares is $\sum (y_i - \bar y)^2$, how much the response varies about its own mean. The residual sum of squares is $\sum e_i^2$, what the line leaves behind. The explained sum of squares is the difference. Their ratio, $R^{2}$, is the explained fraction or coefficient of determination.
Write $\bar y$ for the response mean. Every observation satisfies
$$y_i - \bar y = (\hat y_i - \bar y) + (y_i - \hat y_i),$$
and squaring and summing makes the cross term vanish — because $\sum e_i = 0$ and $\sum x_i e_i = 0$, the two constraints of the last lesson. So
$$\underbrace{\sum (y_i - \bar y)^2}_{SST} = \underbrace{\sum (\hat y_i - \bar y)^2}_{SSR} + \underbrace{\sum e_i^{2}}_{SSE},$$
an exact split of the total variation into the part the line accounts for and the part it leaves. The explained fraction is
$$R^{2} = \frac{SSR}{SST} = 1 - \frac{SSE}{SST}.$$
Because the fitted line does at least as well as the flat line at $\bar y$, $SSE \le SST$, so $R^{2}$ lies between $0$ and $1$. For simple linear regression it is exactly the square of the correlation between $x$ and $y$, which is where the symbol comes from.
What it measures is how much of the variation the line accounts for in these data. Three things it does not measure are worth naming. It does not say the model is the right shape — a pronounced bend can still give a high figure, and only the residual plot shows it. It does not say the relationship is causal: A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that. And it is not a count of points on the line.
It is also not comparable across studies. A fraction of $0.3$ in a field where responses are noisy may be a real finding, and one of $0.95$ in a physical measurement may be disappointing.
Another way: picture
Two horizontal lines drawn on a scatter: the flat line at the response mean, and the fitted line. The total sum of squares is the squared distance of every point from the flat line; the residual sum is the squared distance from the fitted one. The explained fraction is how much of the first was removed by tilting.
Another way: steps
| Field | Typical explained fraction | What it means there |
|---|---|---|
| Physics laboratory | 0.999 | Anything lower suggests a fault |
| Engineering calibration | 0.95 | Routine; expected |
| Agricultural yield | 0.6 | A strong result |
| Human behaviour | 0.2 | Often a real and publishable finding |
The number is the same object in every row and means something different in each. It compares a line against the flat line at the response mean, and how much room there is between those two depends entirely on how noisy the response is — which is a fact about the subject, not about the model.
Reading it as a diagnostic. It says how much variation the line removed, not whether a line was the right thing to fit. A curve can give a high figure and a residual plot that condemns the model.
Reading it as evidence of cause. A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that.
Reading it as a proportion of points. It is a ratio of sums of squares. No observation need lie on the line at all.
Comparing it between studies of different responses. The denominator is the response's own variability, so the two figures are measured against different yardsticks.
Adding explanatory variables to raise it. It can never fall when a variable is added, however useless the variable, which is why it cannot be used to choose between models of different sizes.
$SST = 400$ and $SSE = 32$.
Total and residual.
$SSR = 400 - 32 = 368$, so $R^{2} = 368/400$.
The identity.
That is $0.92$: the line accounts for ninety-two per cent of the variation in these data.
A description, carefully worded.
Points lie on a smooth curve rising then levelling off; a straight line is fitted.
The wrong shape entirely.
The line still removes most of the variation, and $R^{2}$ comes to $0.93$.
Looks excellent.
The residual plot shows a clear arch, and predictions at the ends of the range are badly wrong.
The figure never reports this.
The unexplained fraction is $50/250$.
Residual over total.
That is $0.2$, so the explained fraction is $1 - 0.2$.
The line accounts for $0.8$ of the variation in these data.
A fitted line leaves a residual sum of squares of $132$, and the total sum of squares of the response is $200$. What fraction of the variation does the line account for?
Answer:
Three fits to the same response, whose total sum of squares is $900$. Give the explained sum of squares and the explained fraction for each.
| Explained sum of squares | Explained fraction | |
|---|---|---|
| Residual sum of squares $720$ | ||
| Residual sum of squares $225$ | ||
| Residual sum of squares $36$ |
The total sum of squares of a response is $400$, and a fitted line leaves $276$ of it unexplained. Give the explained sum of squares, and the explained fraction.
Explained sum of squares: g. Explained fraction: f.
A regression on a response whose total sum of squares is $600$ leaves $120$ unexplained. Match each quantity to what it measures.
| How much the response varies about its own mean, before any line | The variation the fitted line leaves behind | The variation the fitted line accounts for | The share of the variation the line accounts for, with no units | |
|---|---|---|---|---|
| $600$ | ||||
| $120$ | ||||
| $480$ | ||||
| $0.8$ |
Four lines are fitted to the same response, whose total sum of squares is $60$. Their residual sums of squares are below. Put them in order of the fraction of the variation they account for, smallest first.
Number the steps in order (write the number in the box):
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A regression of one measured quantity on another, over $60$ observations, accounts for $90\%$ of the variation in the response. What does that entitle anyone to say?
You can compute the explained fraction from two sums of squares and say what each sum measures. Say in your own words why a high explained fraction is not evidence that a straight line was the right model.
9. Your turn: $SST = 250$ and $SSE = 50$, step 3