Back to the on-screen lesson ·
The line built from five summary statistics, what its slope means in context, residuals and the plot that judges the fit, and the range outside which it says nothing.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you can build the least-squares line from five summary statistics alone, interpret its slope as a rate in the units of the data, and compute a residual as observed minus predicted. You can also say what the line does not license: a cause, and a prediction outside the range of $x$ that was actually measured.
You can write the equation of a line from a slope and an intercept, and you have just met $r$, $s_x$ and $s_y$. Those five numbers are enough to build the one line that fits a cloud of points better than any other.
$\hat{y}$: the value the line predicts, said 'y-hat' — never an observed value.
Residual: observed minus predicted, $y - \hat{y}$.
Least-squares line: the line making the total of the squared residuals as small as possible.
Residual plot: the residuals plotted against $x$; it should look like nothing at all.
Extrapolation: using the line outside the range of $x$ that was observed.
The least-squares regression line $\hat{y} = a + bx$ is the line for which the total of the squared residuals is as small as possible. Two facts define it, and both are worth remembering because they let you build it from summary statistics alone:
$$b = r\,\frac{s_y}{s_x}, \qquad \bar{y} = a + b\bar{x}$$
The first says the slope is the correlation, rescaled from standardised units back into the units of the data. The second says the line always passes through the point of averages, which gives $a = \bar{y} - b\bar{x}$.
Interpreting the slope in context is the part that carries meaning: it is the predicted change in $y$ for a one-unit increase in $x$ — an association, not an intervention. The intercept is the prediction at $x = 0$, which is often outside the data and sometimes meaningless.
Two checks stand between a fitted line and a claim. The residual plot should show no pattern; a curve or a smile in it means a straight line was the wrong shape, whatever $r^2$ said. And the line describes the range of $x$ that was observed — outside it, the line is an assumption, which is how a regression comes to predict a negative price.
Another way: picture
Picture a vertical line from every point down or up to the fitted line. Square each of those lengths into an actual square, and the least-squares line is the one position for which the total area of all those squares is smallest. Squares, not distances — which is why one far-off point counts for so much.
Another way: steps
| What the residual plot shows | What it means |
|---|---|
| A shapeless band about zero | A line was a reasonable shape |
| A clear curve or smile | The relationship is curved; refit |
| A funnel, widening to the right | The spread grows with $x$; the usual intervals are unsafe |
| One point far from the rest | Influential; check it and refit without it |
The residual plot is the only one of the four checks that costs nothing and is skipped most often.
"The residual is predicted minus observed." It is observed minus predicted. Get it backwards and every sign in the analysis is inverted, including the reading of the residual plot.
"A slope is a value." A slope is a rate: the change in $\hat{y}$ per one unit of $x$. "Rent is $\pounds12$" is not what a slope of $12$ says.
"A high $r^2$ means the model is right." $r^2$ measures how much variation the line accounts for, not whether a line was the right shape. A curved relationship can produce a respectable $r^2$ and a residual plot that is obviously a smile.
$\bar{x} = 20$, $\bar{y} = 110$, $s_x = 4$, $s_y = 20$, $r = 0.8$.
Five numbers are enough.
$b = 0.8 \times \frac{20}{4} = 4$.
Correlation, rescaled into the data's units.
$a = 110 - 4 \times 20 = 30$, so $\hat{y} = 30 + 4x$.
Through the point of averages.
At $x = 28$: $\hat{y} = 30 + 4 \times 28 = 142$.
Substitute.
The observed value there was $146$, so the residual is $146 - 142 = 4$.
Observed minus predicted.
Positive, so the point sits above the line — the line under-predicted that case.
The sign is the reading.
The prediction at $x = 28$ is $142$.
The residual is $138 - 142 = -4$: below the line.
If the residuals at small $x$ were all positive and those at large $x$ all negative, that pattern — not the size of any one residual — would be the reason to abandon the straight line.
The least-squares line is $\hat{y} = 12 + 1x$. Predict $y$ when $x = 18$.
answer
The line $\hat{y} = 45 + 2x$ predicts a value at $x = 35$, and the observed value there is $117$. What is the residual?
answer
A data set has $\bar{x} = 25$, $\bar{y} = 195$, $s_x = 4$, $s_y = 40$ and $r = 0.6$. Give the slope and then the intercept of the least-squares line.
b = b, a = a
For flats, $\hat{y} = 12 + 3x$, where $x$ is the floor area in square metres and $\hat{y}$ is the monthly rent in pounds. What does the slope $3$ mean?
A shop fits $\hat{y} = 30 + 3x$ to weekly sales $y$ against advertising spend $x$. In the week it spent $28$ it took $117$, and the manager says the campaign 'beat the model'. By how much did that week exceed the prediction?
answer
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
The line $\hat{y} = 30 - 12x$ was fitted to cars aged 1 to 8 years, with $\hat{y}$ the price in hundreds of pounds. Used at $x = 28$ it predicts a negative price. What does that show?
You can fit, use and judge a least-squares line. Without looking: what is the slope in terms of $r$, and which way round is a residual?
9. Your turn: $\hat{y} = 30 + 4x$, observed $y = 138$ at $x = 28$, step 3