Back to the on-screen lesson ·

Hypothesis tests

Null and alternative hypotheses, test statistics as counts of standard errors, $p$-values and what they do not mean, the two kinds of error, power, and the link between a test and a confidence interval.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you can turn a claim into a pair of hypotheses about a parameter, compute a test statistic and its $p$-value, and write a conclusion in context at a threshold chosen before the data. You can also name and distinguish the two kinds of error, explain why lowering one rate raises the other, and say exactly what a $p$-value is — and the three things it is regularly mistaken for.

2. What you bring to this

You can build a confidence interval: estimate, standard error, critical value. A hypothesis test uses the same three pieces to answer a different question — not 'which values are plausible?' but 'is this particular value still plausible?'

3. Words you will need

$H_0$, the null hypothesis: the claim taken at face value, always an equality.

$H_a$, the alternative: what would contradict it, one-sided or two-sided.

Test statistic: how many standard errors the estimate sits from $H_0$.

$p$-value: the probability, assuming $H_0$, of a result at least this extreme.

$\alpha$: the threshold chosen before the data, and the long-run rate of false alarms.

Power: the chance of detecting a real effect.

4. Assume the claim, then see how surprising the data is

A test has one logical move, and it is worth stating plainly because everything else follows from it: assume the null hypothesis, work out how surprising the observed data would be under that assumption, and if it is surprising enough, abandon the assumption.

The steps:

  1. Hypotheses. $H_0$ states the claim as an equality about a parameter; $H_a$ states what would contradict it. Both concern the population; neither concerns the sample.
  2. Conditions. Random sampling, independence, and a sample large enough for the sampling distribution to be approximately normal.
  3. Test statistic. $z = \dfrac{\bar{x} - \mu_0}{\sigma/\sqrt{n}}$, or $t$ with $s$ in place of $\sigma$. It counts standard errors.
  4. $p$-value. The tail area beyond that statistic — both tails if $H_a$ is two-sided.
  5. Conclusion in context. If $p < \alpha$, reject $H_0$; otherwise fail to reject. Quote the $p$-value either way.

Two mistakes are possible and both have names. A type I error rejects a true $H_0$ — a false alarm, at long-run rate $\alpha$. A type II error fails to reject a false $H_0$ — a missed effect, at rate $\beta$. They trade against each other: lowering $\alpha$ to make false alarms rarer makes missed effects commoner. Only more data improves both, by raising the power, $1 - \beta$.

A test and a confidence interval are two views of the same calculation: a two-sided test at $\alpha$ rejects exactly those values that fall outside the corresponding $1 - \alpha$ interval. The interval is usually the more useful report, because it shows the size of the effect and not only whether it cleared a threshold.

Another way: picture

Picture the sampling distribution that $H_0$ predicts, with the observed statistic marked on it. The $p$-value is the area in the tail beyond that mark. Far out in a thin tail, the data is hard to reconcile with $H_0$; near the middle, it is exactly what $H_0$ expected — and the mark itself never says which.

Another way: steps

  1. Write $H_0$ and $H_a$ about the parameter.
  2. Check the conditions before computing anything.
  3. Compute the test statistic: gap over standard error.
  4. Find the $p$-value, doubling it if $H_a$ is two-sided.
  5. Compare with $\alpha$ and write the conclusion in context, with the $p$-value and the effect size.

5. The four outcomes

$H_0$ true$H_0$ false
Reject $H_0$Type I error (rate $\alpha$)Correct — power ($1 - \beta$)
Fail to rejectCorrectType II error (rate $\beta$)

$\alpha$ is chosen before the data. It has to be, because a threshold picked after seeing the result is not a threshold — it is a description of the result.

6. Three things that trip people up

"The $p$-value is the probability that $H_0$ is true." It is computed by assuming $H_0$ is true and asking how surprising the data would then be. It cannot also be a probability about the assumption it started from.

"Failing to reject means accepting." It means the evidence was not strong enough. A study too small to detect anything fails to reject every time, and has shown nothing.

"Significant means large, or important." It means unlikely under $H_0$. With a large enough sample a difference too small to care about becomes significant; with a small enough sample an enormous one does not. Always read the effect size alongside the $p$-value.

7. Testing a claimed mean of 50

  1. $H_0: \mu = 50$, $H_a: \mu \ne 50$; $\sigma = 12$, $n = 36$, $\bar{x} = 54$.

    Two-sided: drift either way matters.

  2. Standard error $= 12/6 = 2$, so $z = (54 - 50)/2 = 2$.

    Two standard errors out.

  3. One tail beyond $z = 2$ holds $0.0228$; doubled, $p = 0.0455$.

    Two-sided doubles it.

  4. $0.0455 < 0.05$, so reject $H_0$: there is evidence the mean is not $50$. Note how close it is — $\alpha = 0.01$ would not have rejected.

    Report the number, not just the verdict.

8. A significant result that does not matter

  1. A sample of $100{,}000$ finds a mean of $50.05$ against $H_0: \mu = 50$, with $\sigma = 12$.

    Enormous sample.

  2. The standard error is $12/316 \approx 0.038$, so $z \approx 1.3$ for a difference of half a tenth.

    The denominator has collapsed.

  3. With a slightly larger sample this becomes significant — and the effect is still $0.05$, which nobody cares about. Significance is not size.

9. Your turn: $H_0: \mu = 100$, $H_a: \mu > 100$, $\sigma = 15$, $n = 25$, $\bar{x} = 106$

  1. Standard error $= 15/5 = 3$, so $z = (106 - 100)/3 = 2$.

  2. One-sided alternative, so $p$ is one tail: $0.0228$.

    No doubling here.

  3. Your turn: work this step out. Its working is at the end of the packet.

    $0.0228 < 0.05$: reject $H_0$. And had the alternative been two-sided, $p$ would have been $0.0455$ — the same data, a different question.

10. Guided practice

A drug is meant to lower blood pressure by 8 points on average. A trial asks whether it achieves less than that. Which pair of hypotheses is right?

11. Guided practice

$H_0$ says the mean is $78$. A sample of $9$ from a population with $\sigma = 18$ has mean $84$. Find the $z$ statistic.

answer

12. Practice

A two-sided test gives $z = 2.5$. What is the $p$-value, to four decimal places?

answer

13. Practice

A two-sided test gives $p = 0.1336$, and the test was run at $\alpha = 0.05$. What is the conclusion?

14. Practice

A drug is meant to lower blood pressure by 8 points on average. A trial asks whether it achieves less than that. The test rejects $H_0$, but $H_0$ was in fact true. Which error is that?

15. Practice

Match each outcome of a test to its name.

Type I errorType II errorPower
$H_0$ true, and rejected
$H_0$ false, and not rejected
$H_0$ false, and rejected

16. Somewhere new

A small trial of a new teaching method reports $p = 0.1$ at $\alpha = 0.05$ and concludes: “the method makes no difference.” If $0.1$ is above $0.05$, what is wrong with that conclusion?

17. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

18. Test question

Two independent samples give standard errors $11$ and $14$. What is the standard error of the difference between their means, as a squared value?

answer

19. What you can do now

You can run and interpret a hypothesis test. Without looking: what is a $p$-value the probability of, and why is failing to reject $H_0$ not the same as showing it true?

Working for the steps left to you

9. Your turn: $H_0: \mu = 100$, $H_a: \mu > 100$, $\sigma = 15$, $n = 25$, $\bar{x} = 106$, step 3