Back to the on-screen lesson ·

P-values and what they do not say

The probability of data at least this extreme under the null, the equivalent reading as the smallest level that would reject, and the five statements it does not support.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will compute a P-value from a tail probability, double it correctly for a two-sided alternative, compare it with a significance level to reach a decision, and read it as the smallest level at which the data would reject. You will also say precisely which statements about a hypothesis a P-value does not support.

2. From a decision to a measure of surprise

The last lesson decided by comparing a statistic with a critical value chosen in advance. That reports a verdict and throws away how close the verdict was. A P-value keeps it, by reporting the probability of data at least this extreme rather than a yes or a no.

3. Words for this lesson

The P-value is the probability, computed under the null hypothesis, of a result at least as extreme as the one observed. At least as extreme is defined by the alternative: one tail for a one-sided alternative, both for a two-sided one. A result is statistically significant at level $\alpha$ when its P-value is below $\alpha$.

4. The probability of data this extreme, if the null were true

Write $T$ for the test statistic and $t$ for the value it took. The P-value is

$$P(T \text{ at least as extreme as } t \mid H_0 \text{ true}),$$

with at least as extreme meaning $T \ge t$ for an upper one-sided alternative, $T \le t$ for a lower one, and $|T| \ge |t|$ for a two-sided one. For a symmetric null distribution the two-sided value is exactly twice the one-sided value, which is why a study that chooses its direction after the fact can halve its P-value by writing one sentence differently.

There is an equivalent definition that is often more useful: the P-value is the smallest significance level at which these data would reject. Every level above it rejects and every level below it does not, so reporting the P-value lets a reader apply their own level without asking for the data.

The conditioning is the whole of it. A P-value is computed on the assumption that the null hypothesis is true. It is therefore a probability about the data, never about the hypothesis. Turning it round into "the probability that the null is true" needs a prior on the hypothesis, which no frequentist test has and no P-value carries.

And the size of the effect is a separate question. Significant means distinguishable from the null value, and nothing more. A large enough sample makes any real effect significant however small it is, so a P-value can never say whether an effect matters. The size of the effect and an interval around it are what answer that, and they should be reported beside it.

Another way: picture

The null distribution as a bell, with the observed statistic marked and everything beyond it shaded. The P-value is the area of the shading. A statistic near the middle shades almost half the curve; one far out shades a sliver — and the sliver is small because the null makes such data unlikely, not because the null is unlikely.

Another way: steps

  1. Compute the test statistic.
  2. Decide which tail or tails the alternative makes extreme.
  3. Find the probability of that region under the null distribution.
  4. Compare with the level, and report the P-value, the estimate and its interval together.

5. The same P-value, two different findings

Two studies of the same question, both reporting $P = 0.03$.

Study AStudy B
Sample size4040 000
Estimated effect8.0 units0.08 units
Standard error3.7 units0.037 units
Statistic2.172.17
P-value0.030.03

The P-values are identical and the findings are not remotely the same. Study A has found an effect of eight units; study B has found one a hundred times smaller, measured a hundred times more precisely. A reader given only the P-value cannot tell them apart, which is the strongest practical argument for reporting the estimate and its interval beside it.

6. Where this goes wrong

Reading it as the probability that the null is true. The P-value conditions on the null; reversing the conditioning needs a prior, and no frequentist test has one.

Reading it as the probability of being wrong to reject. That rate is the significance level, fixed in advance, and it is a property of the procedure.

Reading a large P-value as evidence for the null. It is a failure to find evidence against it, which is what a test with no power also produces.

Reading a small P-value as a large effect. Precision alone produces small P-values, and a large enough study makes any non-zero effect significant.

Reading $0.049$ and $0.051$ as different in kind. The threshold is a convention. Two results either side of it are almost the same result.

7. Computing one, and reading it

  1. A two-sided test gives $z = 2.2$; the upper tail beyond $2.2$ has probability $0.014$.

    One tail.

  2. Symmetry doubles it: $P = 0.028$.

    Both tails, because the alternative is two-sided.

  3. Below $0.05$, so reject at that level — and $0.028$ is also the smallest level that would.

    The two definitions agree.

8. A large P-value, honestly reported

  1. A study of $12$ patients gives $P = 0.4$ for a treatment effect.

    Nowhere near any conventional level.

  2. The correct report is that the data are compatible with no effect — not that there is no effect.

    Failing to reject is not accepting.

  3. The confidence interval, which is wide, is what shows how little the study could have ruled out.

    The interval carries the information the P-value hides.

9. Your turn: a one-sided P-value of $0.02$, with the study planned at $\alpha = 0.05$ and a two-sided alternative

  1. The alternative is two-sided, so both tails count.

    The reported value counts one.

  2. The two-sided P-value is $2 \times 0.02 = 0.04$.

    Symmetry doubles it.

  3. Your turn: work this step out. Its working is at the end of the packet.

    That is still below $0.05$, so the null is rejected — but by less than the one-sided figure suggested.

10. Guided practice

Is this a correct description of a P-value: the probability that the result would repeat in a fresh sample?

11. Guided practice

Three studies report one-sided P-values of $0.05$, $0.025$ and $0.005$. Give the two-sided P-value for each, for a symmetric null distribution.

Two-sided P-value
One-sided $0.05$
One-sided $0.025$
One-sided $0.005$

12. Practice

A statistic falls in the upper tail, and the probability of a value at least that large under the null hypothesis is $0.03$. Give the one-sided P-value and the two-sided P-value, the null distribution being symmetric.

One-sided: a. Two-sided: b.

13. Practice

Four studies report the P-values below, and each was planned at the $5\%$ significance level. Match each to the decision it settles.

Reject, with strong evidence against the nullReject, but only justFail to reject, but only justFail to reject; the data are entirely ordinary under the null
A P-value of $0.0004$
A P-value of $0.04$
A P-value of $0.06$
A P-value of $0.8$

14. Somewhere new

A study of $4$ observations estimates an effect of $6$ units, with a standard error of $6$, giving a statistic of $1$. A second study finds exactly the same effect of $6$ units from $16$ observations. Put the marker at the second study's test statistic.

0 |——————————| 20

Mark the position with a cross, then write the value:

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A study reports a P-value of $9/100$. What is the smallest significance level at which these data would have rejected the null hypothesis?

Answer:

17. What you can do now

You can move between one-sided and two-sided P-values, turn one into a decision, and name what a P-value is not. Say in your own words why a small P-value is not the same thing as a large effect.

Working for the steps left to you

9. Your turn: a one-sided P-value of $0.02$, with the study planned at $\alpha = 0.05$ and a two-sided alternative, step 3