Back to the on-screen lesson ·

t tests and confidence intervals

Test a coefficient against zero or any other value, build a 95 percent interval, and read both without confusing significance with importance.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute a t statistic against any null value, build a 95 percent confidence interval, and state what each does and does not show.

2. What you already have

From the last lesson you can compute a slope's standard error and you know what it measures: the typical distance between a sample's slope and the population slope, if the estimator is unbiased. From statistics you have tested a hypothesis about a mean and built a confidence interval for one.

Testing a regression coefficient works the same way. The estimate replaces the sample mean, its standard error replaces $s/\sqrt{n}$, and every regression table you will ever read reports exactly the numbers this lesson uses.

3. Terms to use precisely

TermWhat it means
Null hypothesisThe value of the coefficient being tested, such as $H_0: \beta_1 = 0$.
t statistic$(\hat\beta - \beta_{H_0})/se(\hat\beta)$: the distance from the null in standard errors.
Critical valueThe cutoff the t statistic is compared with; $1.96$ for a two-sided 5 percent test in large samples.
Significance levelThe chance of rejecting a true null that the test accepts, usually 5 percent.
p-valueThe chance, if the null were true, of a t statistic at least as far from zero as the one observed.
Confidence intervalEstimate plus or minus the critical value times the standard error.
Statistical significanceRejecting the null at the chosen level; not the same as importance.

4. Distance from the null, measured in standard errors

Under the assumptions of the last lesson plus normally distributed errors — or, without normality, in a large sample — the standardized estimate has a t distribution:

$$t = \frac{\hat\beta_1 - \beta_{1,H_0}}{se(\hat\beta_1)}.$$

If the null is true, $t$ is usually small: it falls between $-1.96$ and $1.96$ about 95 percent of the time in large samples. A value outside that range would be surprising if the null held, so we reject the null at the 5 percent level.

The confidence interval turns the test around. The 95 percent interval $\hat\beta_1 \pm 1.96\,se(\hat\beta_1)$ contains exactly the null values that the test would not reject. So one interval answers every possible test at once: zero is rejected if zero lies outside it, minus one is rejected if minus one lies outside it, and so on.

Another way: action

Take a slope of $0.8$ with standard error $0.25$. It sits $3.2$ standard errors above zero. Picture a bell curve centered at zero: $3.2$ is far out in its tail, so zero is not a plausible value.

Another way: steps

  1. State the null value $\beta_{H_0}$ the question asks about.
  2. Subtract it from the estimate.
  3. Divide by the standard error to get $t$.
  4. Compare $|t|$ with the critical value, $1.96$ at 5 percent.
  5. Build the interval $\hat\beta \pm 1.96\,se$ and check it agrees with the test.

5. Why the interval and the test agree

The test does not reject $\beta_{H_0}$ when $-1.96 \le (\hat\beta - \beta_{H_0})/se \le 1.96$. Multiply through by the standard error and rearrange: that is the same as $\hat\beta - 1.96\,se \le \beta_{H_0} \le \hat\beta + 1.96\,se$.

So the 95 percent confidence interval is precisely the set of null values that a 5 percent test would not reject. That is why economists increasingly report intervals instead of stars: an interval shows at a glance both whether zero is plausible and which effect sizes are plausible, and the second is usually the more useful fact.

6. What a confidence interval does and does not mean

The $95$ percent refers to the method, not to one interval. If you drew many samples and built an interval from each, about 95 percent of those intervals would contain the true coefficient. Any particular interval either contains it or does not.

So it is not correct to say there is a 95 percent chance the true slope lies in $[0.31, 1.29]$. It is fair to say the data are consistent with every value in that range and inconsistent, at the 5 percent level, with values outside it. And the interval, like the standard error it is built from, accounts only for sampling variation. If the estimate is biased, the interval is centered in the wrong place, and 95 percent coverage is lost.

7. Significant is not the same as important

A coefficient is statistically significant when its interval excludes zero. That says the effect is unlikely to be exactly zero. It does not say the effect is large enough to matter.

With a million observations, a return to schooling of $0.001$ — a tenth of a percent per year — can be highly significant and economically trivial. With fifty observations, an estimated effect of twenty percent can be insignificant because its interval runs from minus five percent to forty-five. The second study has not shown the effect is zero; it has shown the data are too few to say.

Always read the size of the coefficient and the width of its interval. The question to ask is which effects are ruled out, not merely whether zero is.

A useful habit is to name, before looking at the results, the smallest effect that would matter for the decision at hand — say a one percent change in employment, or a two point gain in test scores. Then compare the interval with that threshold. An interval entirely below it says the effect, whatever it is, is too small to matter; one entirely above it says the effect matters; one that straddles it says more data are needed before the decision can safely rest on this one study alone, however significant it looks.

8. Testing against values other than zero

Zero is the default null in most software, but economic questions often concern other values. Is demand unit elastic, $\beta = -1$? Is a stock's beta equal to one, moving exactly with the market? Is the return to a year of college equal to the return to a year of high school?

The recipe is the same: subtract the null value, divide by the standard error. Forgetting to subtract is a common error: the $t$ printed in a regression table is always against zero, and dividing the estimate by its standard error answers a question nobody asked when the question was about minus one.

9. Why 1.96, and when it is not

The number $1.96$ comes from the standard normal distribution: $95$ percent of its area lies within $1.96$ of zero. In large samples the t statistic is close to normal whatever the distribution of the errors, by the central limit theorem, so $1.96$ is the right cutoff for a two-sided test at the $5$ percent level.

In small samples with normally distributed errors the exact distribution is Student's t with $n - 2$ degrees of freedom, whose tails are fatter. With $10$ degrees of freedom the cutoff is about $2.23$; with $30$, about $2.04$; by $120$ it is $1.98$. Using $1.96$ with ten observations rejects too often. That is why software reports t-based p-values, and why a rough rule of "about two standard errors" is safe for most samples economists work with.

Other levels use other cutoffs. A $10$ percent test uses $1.645$, a $1$ percent test $2.576$. The stricter the test, the wider the matching interval: a $99$ percent interval is about a third wider than a $95$ percent one.

10. Many tests at once

If you test twenty coefficients that are all truly zero, each at the $5$ percent level, you should expect about one of them to look significant by chance. A table with a single star among twenty rows is what pure noise looks like.

The problem is worse when researchers try many specifications and report the one that works — a practice called p-hacking. It is one reason economists now value pre-registered analysis plans, which commit to the tests before the data are seen, and why a surprising result that appears in only one specification deserves skepticism. When several related hypotheses are tested together, the joint test of a later lesson is the right tool, and it controls the overall chance of a false alarm rather than the chance for each coefficient separately.

11. Working a test, step by step

A demand regression estimates a price elasticity of $-1.3$ with standard error $0.15$. Is demand unit elastic?

  1. Null. $H_0: \beta = -1$.
  2. Distance. $-1.3 - (-1) = -0.3$.
  3. t statistic. $-0.3/0.15 = -2$.
  4. Compare. $|-2| = 2 > 1.96$: reject at the 5 percent level.
  5. Interpret. Demand is estimated to be elastic: a price rise would reduce revenue. The interval $-1.3 \pm 0.294$, from about $-1.59$ to $-1.01$, just excludes $-1$, as the test said it would.

12. How to check a test answer

Three checks.

  1. Did you subtract the right null? The table's $t$ is against zero; your question may not be.
  2. Do the test and interval agree? If $|t| > 1.96$, the null value must lie outside your interval. If they disagree, one calculation is wrong.
  3. Is the interval symmetric? The estimate must sit exactly in its middle.

Then read the interval in words: "the data are consistent with effects between this and that." If that range includes both trivial and large effects, say so.

13. In the world: what does the minimum wage do to teen jobs?

Few questions in economics have been studied as intensively as the effect of the minimum wage on employment. Many studies report an elasticity: the percent change in teen employment for a percent change in the minimum wage. Surveys of this literature find estimates clustered between about $-0.3$ and $0$, with many individual studies unable to reject zero.

Suppose a study finds that a $10$ percent increase changes teen employment by $-0.4$ percent, with a standard error of $0.1$. The $95$ percent interval runs from about $-0.6$ to $-0.2$ percent. Zero is rejected, but so is any large job loss: the data rule out, say, a $3$ percent drop. A second study finding $0.35$ with a standard error of $0.2$ cannot reject zero, but its interval, $-0.04$ to $0.74$, also rules out large losses.

Policy debates often turn on the phrase "no significant effect." The useful question is what the intervals rule out. Two insignificant studies with narrow intervals near zero are strong evidence of small effects; one insignificant study with an interval from $-5$ to $+5$ percent is evidence of very little. Reading the intervals, not the stars, is what lets the two kinds of study be told apart, and it is how careful reviews of this literature are written.

14. Failing to reject is not evidence of no effect

The most common mistake is reading an insignificant coefficient as proof that the effect is zero. It means only that zero is among the plausible values; the interval may include large effects too.

A second mistake is reading the $95$ percent as a probability about one interval. It describes how often the method works across samples.

A third is testing every question against zero. When the question is whether an elasticity is minus one or a beta is one, subtract that value before dividing.

15. A t test against zero

  1. State the null hypothesis.

    $H_0: \beta = 0$

    No relationship.

  2. Read the estimate and error.

    $\hat\beta = 1.5, \ se = 0.5$

    From the table.

  3. Divide the two numbers.

    $t = 1.5/0.5 = 3$

    Three standard errors from zero.

  4. Compare with the critical value.

    $3 > 1.96$

    Beyond the 5 percent cutoff.

  5. State the conclusion.

    $\text{reject } H_0$

    Zero is implausible.

16. A 95 percent confidence interval

  1. Read the estimate.

    $\hat\beta = -0.4$

    A negative slope.

  2. Read the standard error.

    $se = 0.1$

    Its sampling uncertainty.

  3. Compute the margin.

    $1.96 \times 0.1 = 0.196$

    About two standard errors.

  4. Find the lower end.

    $-0.4 - 0.196 = -0.596$

    Subtract the margin.

  5. Find the upper end.

    $-0.4 + 0.196 = -0.204$

    Add the margin.

  6. Check against zero.

    $0 \notin [-0.596, -0.204]$

    So zero is rejected, matching $t = -4$.

17. An insignificant but uninformative estimate

  1. Read the estimate.

    $\hat\beta = 3$

    A large point estimate.

  2. Read the standard error.

    $se = 2$

    Very noisy.

  3. Compute the t statistic.

    $t = 3/2 = 1.5$

    Less than 1.96.

  4. Compute the margin.

    $1.96 \times 2 = 3.92$

    Wide.

  5. Build the interval.

    $[-0.92, \ 6.92]$

    Includes zero and large values.

  6. State the test result.

    $\text{do not reject } H_0$

    Zero is plausible.

  7. State what the data allow.

    $\text{effects from about } -1 \text{ to } 7$

    Not evidence of no effect.

18. Your turn: estimate 0.8, standard error 0.25.

  1. Compute the t statistic.

    $0.8/0.25 = 3.2$

    Against zero.

  2. Compute the margin.

    $1.96 \times 0.25 = 0.49$

    About two standard errors.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Build the interval.

19. Guided practice

A regression coefficient is $0.15$ with a standard error of $0.03$. What is the t statistic for the null hypothesis that the coefficient is zero?

Answer: t statistic

20. Guided practice

Complete the worked solution: a return-to-schooling estimate is $0.03$ in a log-wage regression, with standard error $0.0043$. Find the margin of error at 95 percent, and both ends of the interval.

  1. Multiply by the critical value.

    $1.96 \times 0.0043 =$ a

    About two standard errors.

  2. Subtract the margin.

    $0.03 - \text{margin} =$ b

    The lower end.

  3. Add the margin.

    $0.03 + \text{margin} =$ c

    The upper end.

21. Guided practice

An estimated coefficient is $0.35$ with a standard error of $0.2$. Using a critical value of $1.96$, what is the lower end of the 95 percent confidence interval?

Answer: lower end

22. Practice

A coefficient is $3$ with standard error $2$. Fill in the t statistic against zero and both ends of the 95 percent confidence interval (critical value $1.96$).

value
t statistic
lower end of the interval
upper end of the interval

23. Practice

A slope is $0.35$ with a 95 percent confidence interval of $[-0.04, 0.74]$. Match each statement to a verdict.

a correct readinga misreading
At the 5% level we cannot reject that the slope is zero.
The data show the slope is zero.
A slope of 0.6 is consistent with the data.
There is a 95% chance the true slope is 0.35.

24. Practice

A log-log demand regression for a streaming service estimates a price elasticity of $-1.3$ with a standard error of $0.15$. Test the null hypothesis that demand is unit elastic ($\beta_1 = -1$). What is the t statistic?

Answer: t statistic

25. Somewhere new

A study of state minimum-wage increases estimates that a 10 percent increase changes teen employment by $0.35$ percent, with a standard error of $0.2$. What is the upper end of the 95 percent confidence interval (critical value $1.96$)?

Answer: percent

26. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

27. Test question

A coefficient is $3$ with standard error $2$. Fill in the t statistic against zero and both ends of the 95 percent confidence interval (critical value $1.96$).

value
t statistic
lower end of the interval
upper end of the interval

28. What you can do now

You can test a coefficient. Explain to someone why an insignificant estimate is not evidence that an effect is zero.

Working for the steps left to you

18. Your turn: estimate 0.8, standard error 0.25., step 3

$[0.31, \ 1.29]$

Zero lies outside, so reject.