Back to the on-screen lesson ·
How fast a family of comparisons grows, how many false positives it expects when nothing is happening, and the order in which several means are properly compared.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will count the pairwise comparisons a set of groups produces, compute how many false positives a family of tests expects when every null is true, apply a Bonferroni adjustment, and carry out a comparison of several means in the order that keeps the stated error rate honest.
A significant F says the group means are not all equal. It does not say which ones differ, and that is almost always the question somebody actually has. Answering it means running several tests, and several tests behave quite differently from one.
A family of tests is a set of hypotheses reported together. The family-wise error rate is the probability of at least one false positive anywhere in the family. A Bonferroni adjustment divides the level by the number of comparisons. A post-hoc comparison is one chosen after the data were seen; a planned comparison is one chosen before.
With $k$ groups there are
$$\binom{k}{2} = \frac{k(k-1)}{2}$$
pairwise comparisons, a count growing as the square of $k$: three groups give three comparisons, six give fifteen, twenty give a hundred and ninety.
Suppose every null is true and each comparison is tested at level $\alpha$. The expected number of rejections is $\binom{k}{2}\alpha$, whether or not the comparisons are independent, because expectations of indicators simply add. For independent tests the probability of at least one rejection is $1 - (1 - \alpha)^{m}$ for $m$ tests, which reaches $0.40$ at $m = 10$ and $0.64$ at $m = 20$ for $\alpha = 0.05$.
Note carefully what is and is not wrong. Every individual test is behaving exactly as designed, rejecting a true null one time in twenty. The error is in the reporting: a rejection somewhere among many is unremarkable, and presenting it as a single planned finding claims a guarantee that was never in force.
Bonferroni is the simplest repair: test each of $m$ comparisons at $\alpha/m$, and the family-wise rate is at most $\alpha$. It is conservative and it always works, whatever the dependence among the tests. Tukey's method is tighter for all pairwise comparisons of equal-sized groups. Both cost power, which is the honest price of a guarantee that covers the whole family.
The order matters as much as the adjustment. The overall F test comes first and buys the right to look; the comparisons come second; the adjustment always. Choosing which pairs to compare after seeing which look largest, without adjusting, is the abuse the whole procedure exists to prevent.
Another way: story
Twenty laboratories each test a coin they know is fair, each at the five per cent level. One of them, on average, reports a significant result. If that laboratory publishes and the other nineteen do not, the literature contains one finding of a biased coin and no trace of the nineteen that found nothing — and every single test was run correctly.
Another way: steps
| Groups | Comparisons | Expected false positives | Chance of at least one |
|---|---|---|---|
| 3 | 3 | 0.15 | 14% |
| 4 | 6 | 0.30 | 26% |
| 6 | 15 | 0.75 | 54% |
| 10 | 45 | 2.25 | 90% |
| 20 | 190 | 9.50 | over 99% |
The last column assumes independence, which pairwise comparisons are not, so it is an approximation — and the direction of the story is not in doubt. By six groups a table of unadjusted pairwise P-values is more likely than not to contain a false finding, and the study that produced it may be perfectly well conducted.
Believing each test being correct makes the family correct. The level governs one test. Nothing in it governs a set of them.
Reading the significance level as a promise of no false positives. It is the rate at which they happen, and over many tests they will.
Reading it as the share of findings that are false. That quantity depends on how many of the nulls were true to begin with, which no test knows.
Choosing comparisons after seeing which look largest, then not adjusting. This is the most damaging version, because the selection itself is a form of testing that leaves no trace in the output.
Running the pairwise tests without the overall test. The overall test is what makes looking at all defensible.
Five groups, so ten pairwise comparisons, with a family-wise rate of $0.05$ wanted.
Count the comparisons first.
Each comparison is tested at $0.05/10 = 0.005$.
Divide the level.
A pair with $P = 0.01$ is therefore not significant here, though it would be on its own.
This is what the guarantee costs.
A real difference that a single test would detect $80\%$ of the time is one of ten comparisons.
Level now $0.005$.
The critical value moves out, and the power against the same difference falls to around $55\%$.
A quarter of the detections lost.
The remedy is fewer comparisons, chosen in advance — not a weaker adjustment.
Planning beats correcting.
The number of comparisons is $\dfrac{7 \times 6}{2}$.
Unordered pairs.
That is $21$.
So $21 \times 0.05 = 1.05$ false positives are expected — more than one, from a study in which nothing is happening.
$5$ group means are to be compared. How many pairwise comparisons would comparing every pair require?
Answer:
A study of $36$ observations per group compares every pair of means at the $5\%$ level, and in truth all the means are equal. For each number of groups, give the number of comparisons and the expected number of false positives.
| Comparisons | False positives expected | |
|---|---|---|
| With $4$ groups | ||
| With $5$ groups | ||
| With $6$ groups |
$9$ groups are compared pair by pair at the $5\%$ level, and in truth every mean is the same. Give the number of comparisons made, and the expected number of them that reject.
Comparisons: c. Expected rejections: e.
Each group holds $16$ observations. Match each number of groups to the number of pairwise comparisons it produces.
| $3$ comparisons | $6$ comparisons | $10$ comparisons | $15$ comparisons | |
|---|---|---|---|---|
| $3$ groups | ||||
| $4$ groups | ||||
| $5$ groups | ||||
| $6$ groups |
$8$ group means are to be compared, and the analyst wants to know which of them differ. Put the steps into the order they should be carried out.
Number the steps in order (write the number in the box):
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
$160$ independent hypotheses are tested at the $5\%$ level, every null is in fact true, and $8$ of the tests reject. What has gone wrong?
You can count the comparisons a family contains, say how many false positives it expects, and order the steps of comparing several means. Say in your own words why every test in a family can be correct while the report of it is not.
9. Your turn: seven groups, every pair compared at the $5\%$ level, all means truly equal, step 3