Back to the on-screen lesson ·

Two kinds of error, and power

The error the significance level controls, the error it does not, and why a test that finds nothing may have had no chance of finding anything.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will name the four outcomes of a test, compute a power from a Type II error rate and back again, say which of the two errors the significance level controls, and explain why power is a curve rather than a number. You will also say what a non-rejection does and does not establish.

2. The error the level does not control

Everything so far has been computed under the null hypothesis: the significance level, the rejection region, the P-value. All of it controls one error. This lesson looks at the other one, which is invisible from inside the null distribution and is where most of the damage in applied work happens.

3. Words for this lesson

A Type I error is rejecting a true null; its rate is the significance level $\alpha$. A Type II error is failing to reject a false null; its rate is $\beta$, and it depends on how false the null is. The power against an alternative is $1 - \beta$: the probability of detecting that alternative. An effect size is the departure from the null, in the quantity's own units or in standard errors.

4. Two errors, and only one of them is controlled

A test makes one of four outcomes, according to the truth and the decision:

Test rejectsTest fails to reject
Null trueType I error, rate $\alpha$correct, rate $1 - \alpha$
Null falsecorrect, rate $1 - \beta$ (power)Type II error, rate $\beta$

Each row adds to one; the two rows do not, because they are conditional on different states of the world. Nobody ever knows which row they are in.

The asymmetry between the two errors is deliberate. $\alpha$ is chosen in advance and holds for every test of that null. $\beta$ cannot be chosen because it is not a single number: it depends on which alternative is named, and it approaches $1 - \alpha$ as the alternative approaches the null. A test has one significance level and a whole power curve.

Power rises with three things. It rises with the effect size, because a larger departure is easier to see. It rises with the sample size, because the standard error shrinks and the same effect becomes more standard errors. And it rises with $\alpha$, because a wider rejection region catches more of everything — which is a trade rather than an improvement, and the only one of the three that costs something elsewhere.

Fail to reject is not accept. A test that does not reject has found the data compatible with the null, which is also what a test with too little power finds when the null is false. Absence of evidence against a hypothesis is not evidence for it, and the honest report names the power the test had.

Another way: picture

Two bells on one axis: the null distribution centred at the hypothesised value and the alternative centred where the truth actually is, with a vertical cut at the critical value. The null's tail beyond the cut is $\alpha$; the alternative's mass before the cut is $\beta$. Sliding the cut trades one for the other. More data narrows both bells and pulls them apart, which shrinks both at once — and it is the only thing that does.

Another way: steps

  1. Fix $\alpha$ and find the rejection region from the null distribution.
  2. Name the alternative you want to detect.
  3. Find the probability of landing in the rejection region under that alternative.
  4. That is the power; one minus it is $\beta$.

5. One test, a whole curve of powers

A one-sided test at $\alpha = 0.05$ with a standard error of $1$, against a range of true effects.

True effectEffect in standard errorsPower
000.05
110.26
220.64
330.91
440.99

The first row is the null itself, where the power equals the significance level: a test rejects a true null exactly $\alpha$ of the time, by construction. The power then climbs, and the useful range is narrow — most of the journey from useless to near-certain happens between one and three standard errors. A study sized for two standard errors will miss the effect it was built to find more than a third of the time.

6. Where this goes wrong

Reading a non-rejection as evidence for the null. Fail to reject is not accept. A test that does not reject has found the data compatible with the null, which is also what a test with too little power finds when the null is false. Absence of evidence against a hypothesis is not evidence for it, and the honest report names the power the test had.

Treating power as a single number for a test. It is a function of the alternative. The power of the test means nothing until an effect size is named.

Raising the significance level to raise power. It works and it buys the power with false positives. If that trade is acceptable it should be argued for, not made quietly.

Computing power after the study, from the observed effect. Observed power is a function of the P-value and adds nothing. What a finished study needs is its confidence interval, which says what it could and could not have ruled out.

7. Power against a named alternative

  1. $H_0: \mu = 100$ against $\mu > 100$, $\sigma = 15$, $n = 36$, $\alpha = 0.05$: reject when $\bar x > 100 + 1.645 \times 2.5 = 104.1$.

    Rejection region from the null.

  2. If the truth is $\mu = 106$, then $\bar x$ has mean $106$ and standard error $2.5$.

    Now work under the alternative.

  3. Power $= P(\bar x > 104.1) = P(Z > -0.76) \approx 0.78$.

    A twenty-two per cent chance of missing it.

8. The same test, a smaller effect

  1. If instead $\mu = 102$, the rejection boundary is unchanged at $104.1$.

    The null did not move.

  2. Power $= P(Z > 0.84) \approx 0.20$.

    Four times in five, the study finds nothing.

  3. Reporting no significant effect from such a study says almost nothing about whether an effect exists.

    This is what low power looks like.

9. Your turn: a test has a Type II error rate of $0.35$ against the effect the study was designed to detect

  1. The power is $1 - 0.35$.

    Power is the complement of $\beta$.

  2. That is $0.65$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    So the study will miss the effect it was built for more than a third of the time, and a non-rejection from it establishes very little.

10. Guided practice

A test has significance level $0.03$, and against the alternative of interest its Type II error rate is $0.13$. Give the probability of each decision, under each state of the world.

Test rejectsTest fails to reject
The null hypothesis is true
The null hypothesis is false

11. Guided practice

Against a particular alternative, a test fails to reject a false null hypothesis $33\%$ of the time. What is its power against that alternative?

Answer:

12. Practice

A test is run at the $6\%$ significance level. Match each combination of truth and decision to its name.

A Type I errorA Type II errorA correct rejection, whose probability is the powerA correct decision that nobody reports
The null is true, and the test rejects
The null is false, and the test fails to reject
The null is false, and the test rejects
The null is true, and the test fails to reject

13. Practice

A test rejects a true null hypothesis $9\%$ of the time, and fails to reject a particular false null $24\%$ of the time. Give the significance level and the power, both as proportions.

Significance level: a. Power: p.

14. Somewhere new

The true effect is $10$ units. At a sample of $36$ the standard error of the estimate is $5$. Give the true effect measured in standard errors, at each of these sample sizes.

Effect in standard errors
A sample of $36$
A sample of $144$
A sample of $900$

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A study of $20$ patients tests whether a treatment changes a mean outcome, and does not reject the null hypothesis of no change. What has it established?

17. What you can do now

You can fill in the truth table of a test, compute power from a Type II error rate, and name the three things that raise power. Say in your own words why failing to reject is not the same as accepting.

Working for the steps left to you

9. Your turn: a test has a Type II error rate of $0.35$ against the effect the study was designed to detect, step 3