Back to the on-screen lesson ·

Inference for the slope

A standard error for the fitted slope, the interval it gives, the test of no linear relationship, and the three questions that test does not answer.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will count the degrees of freedom a fitted line leaves, build a confidence interval for the true slope, compute the statistic that tests a slope of zero, and read a regression output table line by line. You will also say exactly what rejecting a zero slope does and does not establish.

2. The slope is an estimate too

The slope computed two lessons ago is a number produced from a sample, so it has a sampling distribution like any other estimate. This lesson gives it a standard error and puts it through the machinery of unit two and unit three without changing any of that machinery.

3. Words for this lesson

The residual variance is $\hat\sigma^{2} = SSE/(n - 2)$, dividing by the degrees of freedom the fit leaves. The standard error of the slope is $\hat\sigma/\sqrt{S_{xx}}$. A regression output table reports, for each coefficient, the estimate, its standard error and their ratio. Extrapolation is using the line outside the range of explanatory values it was fitted from.

4. A standard error for a slope, and what follows from it

Under the model $Y_i = \beta_0 + \beta_1 x_i + \varepsilon_i$ with independent $\varepsilon_i$ of mean zero and constant variance $\sigma^{2}$,

$$\hat\sigma^{2} = \frac{SSE}{n - 2}, \qquad \text{SE}(\hat\beta_1) = \frac{\hat\sigma}{\sqrt{S_{xx}}},$$

and for normal errors

$$\frac{\hat\beta_1 - \beta_1}{\text{SE}(\hat\beta_1)} \sim t_{n-2}.$$

The divisor is $n - 2$ because the fit estimated two things. The $S_{xx}$ underneath the standard error is the design's contribution: explanatory values spread widely give a small standard error, and values crowded together give a large one from the same number of observations.

Everything follows from that pivot. A $95\%$ interval is $\hat\beta_1 \pm t_{n-2,\,0.025}\,\text{SE}$. The test of $H_0: \beta_1 = 0$ — no linear relationship — uses the statistic $\hat\beta_1/\text{SE}$, and rejecting is exactly the interval excluding zero.

What the test settles is narrow and worth stating exactly. It asks whether these data are compatible with a slope of zero. It does not ask whether the line is the right shape, which the residual plot answers. It does not ask whether the slope matters, which the interval and the subject answer. And it certainly does not establish cause: A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that.

Another way: picture

The same scatter fitted twice, once with the explanatory values bunched in the middle of the range and once with them spread to both ends. The second line is pinned down at both ends and barely wobbles; the first pivots freely about the point of means. Same number of points, same response noise, very different standard errors.

Another way: steps

  1. Compute the residual sum of squares and divide by $n - 2$.
  2. Take the square root and divide by $\sqrt{S_{xx}}$ for the standard error of the slope.
  3. Build the interval, or divide the slope by the standard error for the statistic.
  4. Report the estimate and the interval, not the verdict alone.

5. A regression output, read line by line

CoefficientEstimateStandard errorStatisticInterval
Intercept17.52.08.7512.9 to 22.1
Slope2.50.3167.911.77 to 3.23

Every row is the same three-step arithmetic: an estimate, its standard error, and their ratio. The intercept's row is usually of no interest, because it tests whether the line passes through the origin and the origin is normally far outside the range of the data. The slope's row is the one the study was run for, and the interval in it says far more than the statistic beside it: a reader learns that slopes below $1.77$ and above $3.23$ are ruled out, which no verdict conveys.

6. Where this goes wrong

Dividing the residual sum of squares by $n$ or $n - 1$. Two things were estimated. Every standard error the fit reports inherits the error.

Reading a significant slope as a cause. A fitted line describes how two measured quantities move together in the data at hand. It says nothing about what would happen if one of them were changed, because nothing in least squares distinguishes a cause from a common cause or from an accident of who was measured. Only the design can do that.

Reading a significant slope as a good fit. A curved relationship produces a significant slope as readily as a straight one.

Reading a non-significant slope as no relationship. With a small $S_{xx}$ or few observations, the interval may contain zero and a great deal else besides. The interval says which.

Using the line outside the range of the data. Nothing in the fit says the relationship continues, and the standard error of a fitted value grows as you move away from the point of means.

7. From sums to an interval

  1. $n = 10$, $SSE = 32$, $S_{xx} = 40$, $\hat\beta_1 = 2.5$.

    Everything the formulas need.

  2. $\hat\sigma^{2} = 32/8 = 4$, so $\hat\sigma = 2$ and the standard error is $2/\sqrt{40} = 0.316$.

    Eight degrees of freedom.

  3. $t_{8,\,0.025} = 2.306$, so the interval is $2.5 \pm 0.729$, that is $1.77$ to $3.23$.

    Zero is well outside.

8. A slope that is significant and trivial

  1. A study of $50\,000$ people fits a slope of $0.004$ with standard error $0.001$.

    An enormous sample.

  2. Statistic $4$, P-value below $0.001$: firmly significant.

    The verdict.

  3. The interval is $0.002$ to $0.006$, which in context is a change nobody could act on.

    The interval is what carries this.

9. Your turn: $n = 12$, slope $6$, standard error of the slope $1.5$

  1. The degrees of freedom are $12 - 2 = 10$.

    Two were spent on the fit.

  2. The statistic is $6/1.5 = 4$.

    Estimate over standard error.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Against a critical value of $2.228$ on $10$ degrees of freedom, the null of zero slope is rejected.

10. Guided practice

A straight line with an intercept is fitted to $27$ points. How many degrees of freedom does the residual variance have?

Answer:

11. Guided practice

A line fitted to $10$ points has slope $7$, and the standard error of that slope is $2$. On $8$ degrees of freedom the two-sided $95\%$ critical value is $2.306$. Give the interval for the true slope, with both endpoints included.

This task has no paper form; do it on a device.

12. Practice

A fit gives an intercept of $33$ with standard error $3$, and a slope of $16$ with standard error $2$. Complete the output table by giving each coefficient's statistic.

Statistic
Intercept: estimate $33$, standard error $3$
Slope: estimate $16$, standard error $2$

13. Practice

A line fitted to $13$ points has slope $27$, with a standard error of $3$. Give the statistic for testing that the true slope is zero, and the degrees of freedom it is compared on.

Statistic: t. Degrees of freedom: d.

14. Somewhere new

Four regressions report the slope estimates and $95\%$ intervals below. Match each to the conclusion it supports.

A clear relationship, measured preciselyNot significant, and consistent with a very large relationship: the study was too small to tellNot significant, and a large relationship has been ruled outSignificant, and small enough that it may not be worth acting on
Slope $50$, interval $45$ to $55$
Slope $50$, interval $-25$ to $125$
Slope $5$, interval $-5$ to $15$
Slope $5$, interval $2.5$ to $7.5$

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A regression on $80$ observations rejects the hypothesis that the true slope is zero. What has the test established?

17. What you can do now

You can build an interval for a slope, test it against zero, and read a regression output. Say in your own words why a significant slope is neither evidence of a good fit nor evidence of a cause.

Working for the steps left to you

9. Your turn: $n = 12$, slope $6$, standard error of the slope $1.5$, step 3