Back to the on-screen lesson ·
A standard error for the fitted slope, the interval it gives, the test of no linear relationship, and the three questions that test does not answer.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will count the degrees of freedom a fitted line leaves, build a confidence interval for the true slope, compute the statistic that tests a slope of zero, and read a regression output table line by line. You will also say exactly what rejecting a zero slope does and does not establish.
The slope computed two lessons ago is a number produced from a sample, so it has a sampling distribution like any other estimate. This lesson gives it a standard error and puts it through the machinery of unit two and unit three without changing any of that machinery.
The residual variance is $\hat\sigma^{2} = SSE/(n - 2)$, dividing by the degrees of freedom the fit leaves. The standard error of the slope is $\hat\sigma/\sqrt{S_{xx}}$. A regression output table reports, for each coefficient, the estimate, its standard error and their ratio. Extrapolation is using the line outside the range of explanatory values it was fitted from.
Under the model $Y_i = \beta_0 + \beta_1 x_i + \varepsilon_i$ with independent $\varepsilon_i$ of mean zero and constant variance $\sigma^{2}$,
$$\hat\sigma^{2} = \frac{SSE}{n - 2}, \qquad \text{SE}(\hat\beta_1) = \frac{\hat\sigma}{\sqrt{S_{xx}}},$$
and for normal errors
$$\frac{\hat\beta_1 - \beta_1}{\text{SE}(\hat\beta_1)} \sim t_{n-2}.$$
The divisor is $n - 2$ because the fit estimated two things. The $S_{xx}$ underneath the standard error is the design's contribution: explanatory values spread widely give a small standard error, and values crowded together give a large one from the same number of observations.
Everything follows from that pivot. A $95\%$ interval is $\hat\beta_1 \pm t_{n-2,\,0.025}\,\text{SE}$. The test of $H_0: \beta_1 = 0$ — no linear relationship — uses the statistic $\hat\beta_1/\text{SE}$, and rejecting is exactly the interval excluding zero.
What the test settles is narrow and worth stating exactly. It asks whether these data are compatible with a slope of zero. It does not ask whether the line is the right shape, which the residual plot answers. It does not ask whether the slope matters, which the interval and the subject answer. And it certainly does not establish cause: A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that.
Another way: picture
The same scatter fitted twice, once with the explanatory values bunched in the middle of the range and once with them spread to both ends. The second line is pinned down at both ends and barely wobbles; the first pivots freely about the point of means. Same number of points, same response noise, very different standard errors.
Another way: steps
| Coefficient | Estimate | Standard error | Statistic | Interval |
|---|---|---|---|---|
| Intercept | 17.5 | 2.0 | 8.75 | 12.9 to 22.1 |
| Slope | 2.5 | 0.316 | 7.91 | 1.77 to 3.23 |
Every row is the same three-step arithmetic: an estimate, its standard error, and their ratio. The intercept's row is usually of no interest, because it tests whether the line passes through the origin and the origin is normally far outside the range of the data. The slope's row is the one the study was run for, and the interval in it says far more than the statistic beside it: a reader learns that slopes below $1.77$ and above $3.23$ are ruled out, which no verdict conveys.
Dividing the residual sum of squares by $n$ or $n - 1$. Two things were estimated. Every standard error the fit reports inherits the error.
Reading a significant slope as a cause. A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that.
Reading a significant slope as a good fit. A curved relationship produces a significant slope as readily as a straight one.
Reading a non-significant slope as no relationship. With a small $S_{xx}$ or few observations, the interval may contain zero and a great deal else besides. The interval says which.
Using the line outside the range of the data. Nothing in the fit says the relationship continues, and the standard error of a fitted value grows as you move away from the point of means.
$n = 10$, $SSE = 32$, $S_{xx} = 40$, $\hat\beta_1 = 2.5$.
Everything the formulas need.
$\hat\sigma^{2} = 32/8 = 4$, so $\hat\sigma = 2$ and the standard error is $2/\sqrt{40} = 0.316$.
Eight degrees of freedom.
$t_{8,\,0.025} = 2.306$, so the interval is $2.5 \pm 0.729$, that is $1.77$ to $3.23$.
Zero is well outside.
A study of $50\,000$ people fits a slope of $0.004$ with standard error $0.001$.
An enormous sample.
Statistic $4$, P-value below $0.001$: firmly significant.
The verdict.
The interval is $0.002$ to $0.006$, which in context is a change nobody could act on.
The interval is what carries this.
The degrees of freedom are $12 - 2 = 10$.
Two were spent on the fit.
The statistic is $6/1.5 = 4$.
Estimate over standard error.
Against a critical value of $2.228$ on $10$ degrees of freedom, the null of zero slope is rejected.
A straight line with an intercept is fitted to $27$ points. How many degrees of freedom does the residual variance have?
Answer:
A line fitted to $10$ points has slope $7$, and the standard error of that slope is $2$. On $8$ degrees of freedom the two-sided $95\%$ critical value is $2.306$. Give the interval for the true slope, with both endpoints included.
This task has no paper form; do it on a device.
A fit gives an intercept of $33$ with standard error $3$, and a slope of $16$ with standard error $2$. Complete the output table by giving each coefficient's statistic.
| Statistic | |
|---|---|
| Intercept: estimate $33$, standard error $3$ | |
| Slope: estimate $16$, standard error $2$ |
A line fitted to $13$ points has slope $27$, with a standard error of $3$. Give the statistic for testing that the true slope is zero, and the degrees of freedom it is compared on.
Statistic: t. Degrees of freedom: d.
Four regressions report the slope estimates and $95\%$ intervals below. Match each to the conclusion it supports.
| A clear relationship, measured precisely | Not significant, and consistent with a very large relationship: the study was too small to tell | Not significant, and a large relationship has been ruled out | Significant, and small enough that it may not be worth acting on | |
|---|---|---|---|---|
| Slope $50$, interval $45$ to $55$ | ||||
| Slope $50$, interval $-25$ to $125$ | ||||
| Slope $5$, interval $-5$ to $15$ | ||||
| Slope $5$, interval $2.5$ to $7.5$ |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A regression on $80$ observations rejects the hypothesis that the true slope is zero. What has the test established?
You can build an interval for a slope, test it against zero, and read a regression output. Say in your own words why a significant slope is neither evidence of a good fit nor evidence of a cause.
9. Your turn: $n = 12$, slope $6$, standard error of the slope $1.5$, step 3