Back to the on-screen lesson ·
The four assumptions behind a fitted line, the plots that examine them, the one that no plot can examine, and the claims a sound model still does not support.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will name the assumptions a regression rests on and the check each one needs, read a residual plot for the wrong shape, an unequal spread and a single influential point, compute the residual spread a fit leaves, and carry out an analysis in the order that puts the diagnostics before the coefficients.
The slope, the explained fraction, the standard errors and the P-values of the last four lessons were all computed under one model. None of them can report that the model is wrong, because each was derived by assuming it. This lesson is where that assumption is examined, and the examination is a picture rather than a number.
Leverage is how far an observation's explanatory value lies from the mean of the explanatory values; a high-leverage point can move the line a long way on its own. An influential point is one whose removal changes the fit noticeably. Heteroscedasticity is error variance that is not constant. Extrapolation is using the line outside the range of explanatory values it was fitted from.
The model $Y_i = \beta_0 + \beta_1 x_i + \varepsilon_i$ carries four assumptions, and they are not equally important or equally checkable.
Linearity — the mean response really is a straight line in $x$. This is the assumption whose failure does the most damage, because the fitted coefficients then estimate nothing in particular. It shows as a curve in the plot of residuals against fitted values.
Constant variance — the spread of the errors is the same all along the line. Its failure leaves the slope roughly right and every standard error, interval and P-value wrong. It shows as a fan in the same plot.
Normality of the errors — needed for the exact $t$ distributions. It matters least: the central limit theorem makes large samples forgiving. It shows in a normal quantile plot of the residuals.
Independence — the errors do not inform one another. Its failure can be severe and no plot of residuals against fitted values will show it. It is checked by asking how the data were collected: measurements in time order, repeated on the same subject, or clustered by school, clinic or household.
Separately from all four, two observations are worth looking for: a point with high leverage, far out in $x$, which can drag the line towards itself, and an influential point whose removal visibly changes the fit. Neither is an error to be deleted. Both are facts about the data that a reader is entitled to know.
Finally, two claims survive every diagnostic and are still unsupported. A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that. And nothing in the fit licenses extrapolation beyond the range of explanatory values observed.
Another way: picture
Four residual plots side by side. The first is a structureless band about zero — the one a sound fit produces. The second arches up and back down: the wrong shape. The third widens from left to right: the wrong error assumption. The fourth is a tidy band with one point far out to the right, sitting almost exactly on zero because the line was dragged over to meet it.
Another way: steps
| Assumption fails | Slope | Standard errors and intervals | Usual repair |
|---|---|---|---|
| Linearity | estimates nothing meaningful | meaningless too | transform, or fit a curve |
| Constant variance | roughly right | wrong, often badly | transform the response, or weight |
| Normality | right | slightly off in small samples | nothing, if the sample is large |
| Independence | right on average | far too small | model the structure explicitly |
The two rows that matter most are the first and the last, and the last is the one no residual plot reports. A study of two hundred measurements taken from twenty patients has twenty independent units and not two hundred, and every interval computed as though it had two hundred is about three times too narrow.
Treating a high explained fraction as a diagnostic. It reports how much variation the line removed, not whether a line was appropriate. A pronounced curve can still give a high figure.
Checking the diagnostics after quoting the coefficients. The plot decides whether the coefficients mean anything, so it comes first.
Plotting residuals against the observed response. The response contains the fitted value, so that plot shows a relationship by construction. Plot against the fitted values or against the explanatory variable.
Deleting an influential point because it is inconvenient. It is a fact about the data. Report the fit with and without it, and say which.
Believing a sound model licenses a causal claim. A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that.
Residuals scattered in an even band about zero, no curve and no fan; the quantile plot roughly straight.
Three assumptions satisfied.
The data were one measurement per independent subject, so the fourth is satisfied by the design.
The one no plot shows.
Only now are the slope and its interval quoted, and only for the range of explanatory values observed.
With the range stated.
The explained fraction is $0.94$ and the slope is strongly significant.
The output looks excellent.
The residual plot arches: negative at both ends, positive in the middle.
The straight line cannot follow the bend.
The coefficients are not reported. A transformation or a curve is fitted instead, and the diagnostics are run again.
The output never said any of this.
The shape is fine — there is no curve — so linearity is not the problem.
Shape first, then spread.
The spread grows, so the constant variance assumption has failed.
That is what a fan means.
The slope is roughly right and every standard error, interval and P-value is wrong; the explained fraction says nothing about it.
A line has been fitted to $62$ observations. Match each assumption of the model to the thing that checks it.
| Residuals against fitted values: look for a curve in the shape | Residuals against fitted values: look for a fan in the spread | A normal quantile plot of the residuals | How the data were collected: order, clustering, repeated measurement | |
|---|---|---|---|---|
| The relationship is linear | ||||
| The error variance is constant | ||||
| The errors are normally distributed | ||||
| The errors are independent of one another |
A regression is fitted and its residuals show residuals negative at both ends and positive in the middle. What has that found?
A line with an intercept is fitted to $11$ observations, and the residual sum of squares is $576$. Fill in the three lines below.
| Value | |
|---|---|
| Degrees of freedom | |
| Residual variance | |
| Residual standard deviation |
A study of $49$ observations is to be analysed by fitting a straight line. Put the steps into the order they should be carried out.
Number the steps in order (write the number in the box):
A regression of a health measure on hours of exercise, over $86$ adults whose exercise ranged from $1$ to $8$ hours a week, gives a slope of $6$ and passes every diagnostic. Four conclusions are drawn. Mark the two the regression does not support.
This task has no paper form; do it on a device.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A line with an intercept is fitted to $34$ observations, leaving a residual sum of squares of $512$. Give the residual variance, and the residual standard deviation.
Residual variance: v. Residual standard deviation: s.
You can match each assumption to its check, read a residual plot, and name the claims a sound fit still does not support. Say in your own words why no plot of residuals against fitted values can report a failure of independence.
9. Your turn: the residuals fan out steadily as the fitted values grow, and the explained fraction is $0.91$, step 3