Back to the on-screen lesson ·
Null and alternative hypotheses, test statistics as counts of standard errors, $p$-values and what they do not mean, the two kinds of error, power, and the link between a test and a confidence interval.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you can turn a claim into a pair of hypotheses about a parameter, compute a test statistic and its $p$-value, and write a conclusion in context at a threshold chosen before the data. You can also name and distinguish the two kinds of error, explain why lowering one rate raises the other, and say exactly what a $p$-value is — and the three things it is regularly mistaken for.
You can build a confidence interval: estimate, standard error, critical value. A hypothesis test uses the same three pieces to answer a different question — not 'which values are plausible?' but 'is this particular value still plausible?'
$H_0$, the null hypothesis: the claim taken at face value, always an equality.
$H_a$, the alternative: what would contradict it, one-sided or two-sided.
Test statistic: how many standard errors the estimate sits from $H_0$.
$p$-value: the probability, assuming $H_0$, of a result at least this extreme.
$\alpha$: the threshold chosen before the data, and the long-run rate of false alarms.
Power: the chance of detecting a real effect.
A test has one logical move, and it is worth stating plainly because everything else follows from it: assume the null hypothesis, work out how surprising the observed data would be under that assumption, and if it is surprising enough, abandon the assumption.
The steps:
Two mistakes are possible and both have names. A type I error rejects a true $H_0$ — a false alarm, at long-run rate $\alpha$. A type II error fails to reject a false $H_0$ — a missed effect, at rate $\beta$. They trade against each other: lowering $\alpha$ to make false alarms rarer makes missed effects commoner. Only more data improves both, by raising the power, $1 - \beta$.
A test and a confidence interval are two views of the same calculation: a two-sided test at $\alpha$ rejects exactly those values that fall outside the corresponding $1 - \alpha$ interval. The interval is usually the more useful report, because it shows the size of the effect and not only whether it cleared a threshold.
Another way: picture
Picture the sampling distribution that $H_0$ predicts, with the observed statistic marked on it. The $p$-value is the area in the tail beyond that mark. Far out in a thin tail, the data is hard to reconcile with $H_0$; near the middle, it is exactly what $H_0$ expected — and the mark itself never says which.
Another way: steps
| $H_0$ true | $H_0$ false | |
|---|---|---|
| Reject $H_0$ | Type I error (rate $\alpha$) | Correct — power ($1 - \beta$) |
| Fail to reject | Correct | Type II error (rate $\beta$) |
$\alpha$ is chosen before the data. It has to be, because a threshold picked after seeing the result is not a threshold — it is a description of the result.
"The $p$-value is the probability that $H_0$ is true." It is computed by assuming $H_0$ is true and asking how surprising the data would then be. It cannot also be a probability about the assumption it started from.
"Failing to reject means accepting." It means the evidence was not strong enough. A study too small to detect anything fails to reject every time, and has shown nothing.
"Significant means large, or important." It means unlikely under $H_0$. With a large enough sample a difference too small to care about becomes significant; with a small enough sample an enormous one does not. Always read the effect size alongside the $p$-value.
$H_0: \mu = 50$, $H_a: \mu \ne 50$; $\sigma = 12$, $n = 36$, $\bar{x} = 54$.
Two-sided: drift either way matters.
Standard error $= 12/6 = 2$, so $z = (54 - 50)/2 = 2$.
Two standard errors out.
One tail beyond $z = 2$ holds $0.0228$; doubled, $p = 0.0455$.
Two-sided doubles it.
$0.0455 < 0.05$, so reject $H_0$: there is evidence the mean is not $50$. Note how close it is — $\alpha = 0.01$ would not have rejected.
Report the number, not just the verdict.
A sample of $100{,}000$ finds a mean of $50.05$ against $H_0: \mu = 50$, with $\sigma = 12$.
Enormous sample.
The standard error is $12/316 \approx 0.038$, so $z \approx 1.3$ for a difference of half a tenth.
The denominator has collapsed.
With a slightly larger sample this becomes significant — and the effect is still $0.05$, which nobody cares about. Significance is not size.
Standard error $= 15/5 = 3$, so $z = (106 - 100)/3 = 2$.
One-sided alternative, so $p$ is one tail: $0.0228$.
No doubling here.
$0.0228 < 0.05$: reject $H_0$. And had the alternative been two-sided, $p$ would have been $0.0455$ — the same data, a different question.
A drug is meant to lower blood pressure by 8 points on average. A trial asks whether it achieves less than that. Which pair of hypotheses is right?
$H_0$ says the mean is $78$. A sample of $9$ from a population with $\sigma = 18$ has mean $84$. Find the $z$ statistic.
answer
A two-sided test gives $z = 2.5$. What is the $p$-value, to four decimal places?
answer
A two-sided test gives $p = 0.1336$, and the test was run at $\alpha = 0.05$. What is the conclusion?
A drug is meant to lower blood pressure by 8 points on average. A trial asks whether it achieves less than that. The test rejects $H_0$, but $H_0$ was in fact true. Which error is that?
Match each outcome of a test to its name.
| Type I error | Type II error | Power | |
|---|---|---|---|
| $H_0$ true, and rejected | |||
| $H_0$ false, and not rejected | |||
| $H_0$ false, and rejected |
A small trial of a new teaching method reports $p = 0.1$ at $\alpha = 0.05$ and concludes: “the method makes no difference.” If $0.1$ is above $0.05$, what is wrong with that conclusion?
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Two independent samples give standard errors $11$ and $14$. What is the standard error of the difference between their means, as a squared value?
answer
You can run and interpret a hypothesis test. Without looking: what is a $p$-value the probability of, and why is failing to reject $H_0$ not the same as showing it true?
9. Your turn: $H_0: \mu = 100$, $H_a: \mu > 100$, $\sigma = 15$, $n = 25$, $\bar{x} = 106$, step 3