Back to the on-screen lesson ·
The error the significance level controls, the error it does not, and why a test that finds nothing may have had no chance of finding anything.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will name the four outcomes of a test, compute a power from a Type II error rate and back again, say which of the two errors the significance level controls, and explain why power is a curve rather than a number. You will also say what a non-rejection does and does not establish.
Everything so far has been computed under the null hypothesis: the significance level, the rejection region, the P-value. All of it controls one error. This lesson looks at the other one, which is invisible from inside the null distribution and is where most of the damage in applied work happens.
A Type I error is rejecting a true null; its rate is the significance level $\alpha$. A Type II error is failing to reject a false null; its rate is $\beta$, and it depends on how false the null is. The power against an alternative is $1 - \beta$: the probability of detecting that alternative. An effect size is the departure from the null, in the quantity's own units or in standard errors.
A test makes one of four outcomes, according to the truth and the decision:
| Test rejects | Test fails to reject | |
|---|---|---|
| Null true | Type I error, rate $\alpha$ | correct, rate $1 - \alpha$ |
| Null false | correct, rate $1 - \beta$ (power) | Type II error, rate $\beta$ |
Each row adds to one; the two rows do not, because they are conditional on different states of the world. Nobody ever knows which row they are in.
The asymmetry between the two errors is deliberate. $\alpha$ is chosen in advance and holds for every test of that null. $\beta$ cannot be chosen because it is not a single number: it depends on which alternative is named, and it approaches $1 - \alpha$ as the alternative approaches the null. A test has one significance level and a whole power curve.
Power rises with three things. It rises with the effect size, because a larger departure is easier to see. It rises with the sample size, because the standard error shrinks and the same effect becomes more standard errors. And it rises with $\alpha$, because a wider rejection region catches more of everything — which is a trade rather than an improvement, and the only one of the three that costs something elsewhere.
Fail to reject is not accept. A test that does not reject has found the data compatible with the null, which is also what a test with too little power finds when the null is false. Absence of evidence against a hypothesis is not evidence for it, and the honest report names the power the test had.
Another way: picture
Two bells on one axis: the null distribution centred at the hypothesised value and the alternative centred where the truth actually is, with a vertical cut at the critical value. The null's tail beyond the cut is $\alpha$; the alternative's mass before the cut is $\beta$. Sliding the cut trades one for the other. More data narrows both bells and pulls them apart, which shrinks both at once — and it is the only thing that does.
Another way: steps
A one-sided test at $\alpha = 0.05$ with a standard error of $1$, against a range of true effects.
| True effect | Effect in standard errors | Power |
|---|---|---|
| 0 | 0 | 0.05 |
| 1 | 1 | 0.26 |
| 2 | 2 | 0.64 |
| 3 | 3 | 0.91 |
| 4 | 4 | 0.99 |
The first row is the null itself, where the power equals the significance level: a test rejects a true null exactly $\alpha$ of the time, by construction. The power then climbs, and the useful range is narrow — most of the journey from useless to near-certain happens between one and three standard errors. A study sized for two standard errors will miss the effect it was built to find more than a third of the time.
Reading a non-rejection as evidence for the null. Fail to reject is not accept. A test that does not reject has found the data compatible with the null, which is also what a test with too little power finds when the null is false. Absence of evidence against a hypothesis is not evidence for it, and the honest report names the power the test had.
Treating power as a single number for a test. It is a function of the alternative. The power of the test means nothing until an effect size is named.
Raising the significance level to raise power. It works and it buys the power with false positives. If that trade is acceptable it should be argued for, not made quietly.
Computing power after the study, from the observed effect. Observed power is a function of the P-value and adds nothing. What a finished study needs is its confidence interval, which says what it could and could not have ruled out.
$H_0: \mu = 100$ against $\mu > 100$, $\sigma = 15$, $n = 36$, $\alpha = 0.05$: reject when $\bar x > 100 + 1.645 \times 2.5 = 104.1$.
Rejection region from the null.
If the truth is $\mu = 106$, then $\bar x$ has mean $106$ and standard error $2.5$.
Now work under the alternative.
Power $= P(\bar x > 104.1) = P(Z > -0.76) \approx 0.78$.
A twenty-two per cent chance of missing it.
If instead $\mu = 102$, the rejection boundary is unchanged at $104.1$.
The null did not move.
Power $= P(Z > 0.84) \approx 0.20$.
Four times in five, the study finds nothing.
Reporting no significant effect from such a study says almost nothing about whether an effect exists.
This is what low power looks like.
The power is $1 - 0.35$.
Power is the complement of $\beta$.
That is $0.65$.
So the study will miss the effect it was built for more than a third of the time, and a non-rejection from it establishes very little.
A test has significance level $0.03$, and against the alternative of interest its Type II error rate is $0.13$. Give the probability of each decision, under each state of the world.
| Test rejects | Test fails to reject | |
|---|---|---|
| The null hypothesis is true | ||
| The null hypothesis is false |
Against a particular alternative, a test fails to reject a false null hypothesis $33\%$ of the time. What is its power against that alternative?
Answer:
A test is run at the $6\%$ significance level. Match each combination of truth and decision to its name.
| A Type I error | A Type II error | A correct rejection, whose probability is the power | A correct decision that nobody reports | |
|---|---|---|---|---|
| The null is true, and the test rejects | ||||
| The null is false, and the test fails to reject | ||||
| The null is false, and the test rejects | ||||
| The null is true, and the test fails to reject |
A test rejects a true null hypothesis $9\%$ of the time, and fails to reject a particular false null $24\%$ of the time. Give the significance level and the power, both as proportions.
Significance level: a. Power: p.
The true effect is $10$ units. At a sample of $36$ the standard error of the estimate is $5$. Give the true effect measured in standard errors, at each of these sample sizes.
| Effect in standard errors | |
|---|---|
| A sample of $36$ | |
| A sample of $144$ | |
| A sample of $900$ |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A study of $20$ patients tests whether a treatment changes a mean outcome, and does not reject the null hypothesis of no change. What has it established?
You can fill in the truth table of a test, compute power from a Type II error rate, and name the three things that raise power. Say in your own words why failing to reject is not the same as accepting.
9. Your turn: a test has a Type II error rate of $0.35$ against the effect the study was designed to detect, step 3