Back to the on-screen lesson ·

Why a good test for a rare thing misleads

A test that is right about nearly everybody still gives mostly false alarms when the thing it looks for is rare.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will take a base rate and a test's error rate, imagine ten thousand people, and fill in the four boxes that follow. From them you will work out how many positive results there are for each real case, and what share of positive results are real, as a whole number of percent. You will be able to say why that share is a different number from the test's accuracy, and why changing how common the condition is changes the answer without changing the test.

2. What you already have

You can build a two-by-two table and name its four boxes: true positive, false positive, false negative, true negative. This lesson changes one number in that table — how many people are in the top row at all — and watches the answer turn upside down.

3. Updating terms

TermWhat it means
Base rateThe prevalence of a specified property in the relevant reference population.
Prior probabilityProbability given background information before incorporating the specified new evidence.
Posterior probabilityProbability after incorporating that evidence under the stated model.
OddsProbability for a hypothesis divided by probability for its alternative.
Likelihood ratioProbability of the evidence under the hypothesis divided by its probability under the alternative.

4. The test's accuracy and the meaning of a positive result are two different numbers

Take a condition that 10 people in 10,000 have, and a test that never misses anybody who has it and is positive for fewer than 2 healthy people in 100. That does not yet tell us how informative a positive result is.

Ten thousand people are checked in this illustrative dataset.

Tests positiveTests negativeTotal
Has it10010
Does not1909,8009,990
Total2009,80010,000

Two hundred people are told they tested positive. Ten of them have it.

Among the positive results in this complete illustrative dataset, 5 percent have the specified condition. That is the conditional proportion for a uniformly selected positive record, not an individual diagnosis. The table records ten true positives and 190 false positives; it does not establish whether the performance is acceptable for any actual use.

The reason is arithmetic, not medicine. The bottom row of the table is enormous, so even a small slice of it — 190 people — swamps the tiny top row. A small share of a huge group beats a large share of a tiny one.

For the comparison table below, use a false-positive rate of exactly 2 percent and perfect sensitivity. At a base rate of 1,000 in 10,000, the expected positives become 1,000 real and 180 false — about 85 percent real. The comparison holds those test rates fixed while changing the base rate.

Another way: picture

Picture ten thousand people in a hall, with a rope across the front penning in the ten who have the condition. The test sends everyone it flags to one side of the room. All ten from behind the rope walk over — and so do 190 people from the vast crowd behind them. Look at that group and you cannot tell them apart, and only one in twenty came from behind the rope.

Another way: steps

To work out what a positive result is worth:

  1. Pick a round number of people to imagine — ten thousand does for anything rare.
  2. Split them by the base rate: how many have it, how many do not.
  3. Run the test down each row separately. Every real case that the test catches; the stated share of the healthy row.
  4. Add the two positive boxes. That is everybody told yes.
  5. Divide the real positives by that total. That is what a positive result means.

5. Hold the rates fixed when comparing base rates

For a controlled numerical comparison, assume an inspection flags every defective component and exactly 2 percent of components without the defect in expectation. Imagine 10,000 components at each of three base rates. With ten defects, the expected true-positive count is ten and the expected false-positive count is 199.8, giving about 4.8 percent true defects among flags. With one hundred defects, the expected counts are one hundred and 198, giving about 33.6 percent. With one thousand defects, they are one thousand and 180, giving about 84.7 percent.

The fractional expected count is not a claim that part of a component can be flagged. It represents an average over repetitions under the stated probability model. An actual batch yields integer counts that can differ from those expectations. The earlier illustrative table used 190 observed false positives among 9,990 noncases, so its observed rate is slightly below 2 percent. Do not silently equate that observed count with the exact-rate comparison.

These calculations hold sensitivity and false-positive rate fixed. In an actual new population those rates may change too, for example because the kinds of defect or the conditions of inspection differ. The calculation isolates the mathematical effect of prevalence; it does not promise that every real-world transfer changes only one number. Whether further inspection or any other action is worthwhile also depends on consequences, costs, and the relevant decision rules.

6. Why the mistake is so easy to make

Ask what proportion of this illustrative dataset is classified correctly? and the answer is 98.1 percent. Ask I tested positive, do I have it? and the answer is 5 percent. Both are about the same test on the same day, and people slide from one to the other without noticing, because English uses the same words for both.

The difference is which group you are inside. How often is the test right is a question about everybody tested. Does my positive result mean anything is a question about the people who tested positive, which is a different and much smaller group, and it is mostly made of healthy people who got an unlucky result.

Saying which group the question is about is the same move as lesson 1. It has simply been dressed in arithmetic.

7. Bayes' rule in odds form

A table makes the groups visible. Bayes' rule expresses the same update using probabilities or odds. Let H mean that a fictional component has a specified defect, and let E mean that an inspection flags it. The prior probability P(H) describes the defect probability before learning this inspection result, given the background information already included. The posterior probability P(H given E) describes it after incorporating the flag. Prior does not mean uninformed guess: a relevant base rate can be one source of prior information.

Probability and odds use different denominators. If the defect probability is one in ten, there is one defective case for nine nondefective cases. The odds for defect are therefore one to nine, or the ratio 1/9. In general odds equal p divided by one minus p. To convert an odds ratio o back into probability, use o divided by one plus o. For odds a to b, the probability is a divided by a plus b. Confusing odds of one to four with probability one in four changes 20 percent into 25 percent.

The likelihood ratio compares how probable the observed evidence is under the two alternatives. For a positive inspection it is P(E given H) divided by P(E given not H): sensitivity divided by false-positive rate. Suppose those probabilities are 0.8 and 0.2. A flag is four times as probable when the specified defect is present, so the likelihood ratio is four. This does not mean the flagged component is four times as likely to be defective as not defective. That latter comparison also needs the prior odds.

The odds form of Bayes' rule is: posterior odds equal prior odds multiplied by the likelihood ratio. With prior odds one to nine and likelihood ratio four, the posterior odds are four to nine. The resulting probability is 4/(4 + 9) = 4/13, about 30.8 percent. The flag increases the probability from ten percent, but the nondefective alternative remains more probable. Evidence can favor a hypothesis without making it more probable than its alternative overall.

A frequency table checks this result. In a modeled batch of one thousand components, one hundred have the defect and nine hundred do not. The expected true-positive count is eighty; the expected false-positive count is 180. Among the 260 expected flags, the defective share is 80/260 = 4/13. The table and odds method agree because they express the same conditional-probability relationship, not because one has supplied extra evidence missing from the other.

For these calculations assume probabilities strictly between zero and one and a nonzero probability for the evidence. The simple odds division needs care at zero denominators. More fundamentally, the rule cannot repair inappropriate input probabilities. A defect rate from a different production process may be a poor prior for this batch, and a test rate estimated under ideal conditions may not describe current operation. Explain which population and conditions your inputs are intended to represent.

8. Evidence strength, uncertainty, and repeated signals

A likelihood ratio greater than one increases the odds of H; a ratio below one decreases them; a ratio of one leaves them unchanged. The direction depends on a comparison between the evidence under H and under its alternative. Evidence being common under H is not enough by itself. If a signal is just as common when H is absent, observing it does not discriminate between these alternatives under the model.

A negative result has its own likelihood ratio. If sensitivity is 80 percent and the false-positive rate is 20 percent, then P(negative given defect) is 20 percent and P(negative given no defect) is 80 percent. The negative-result likelihood ratio is therefore one quarter. With prior odds one to nine, the updated odds are one to 36, and the probability is one in 37, about 2.7 percent. A negative result reduces the probability in this model without proving absence.

Do not multiply the same likelihood ratio twice merely because a result appears in two reports. If two reports describe the same inspection, the second report is not a second independent observation. Even two separate inspections can share an error source such as the same contaminated sensor. For sequential updating, the second likelihood ratio must compare its probabilities conditional on the evidence already received. Reusing the original ratio is justified only when the required conditional independence or equivalent model assumptions hold.

Uncertainty in the inputs should remain visible in the conclusion. If the base rate or test rates are estimated from limited observations, a neatly calculated posterior is conditional on those estimates. Exploring several plausible input values can show how much the conclusion depends on them. This lesson uses fixed fictional numbers to teach the reasoning, so its exact answers belong to the stated models. Real evaluation also needs evidence about measurement, sampling, stability, and the costs of errors.

Finally keep a probability statement separate from an action rule. A 30.8 percent defect probability might justify another inspection in one setting and a different response in another. Neither Bayes' rule nor a large accuracy percentage sets those costs or priorities. The reasoning contribution is to state what the evidence implies under explicit assumptions, so that a decision is not made by mistaking a flag for certainty or by dismissing informative evidence merely because it falls short of certainty.

9. A fictional component flag, checked two ways

A manufacturer models a batch of 1,000 components using a ten-percent defect rate. Its inspection flags 80 percent of defective components and 20 percent of components without the defect. Treat these as the stated model probabilities, and suppose the result available about a selected component is a positive flag. The question is what the model predicts among flags, not whether the device is suitable for a particular industrial use.

The frequency route starts with one hundred defective and nine hundred nondefective components in expectation. Eighty of the defective group are flagged, and 180 of the nondefective group are flagged. There are 260 flags in total. The expected defective share among flags is 80 divided by 260, about 30.8 percent. The other twenty defective components are missed, while 720 nondefective components receive negative results.

The odds route starts with prior odds one to nine. A flag is 0.8 divided by 0.2, or four times as likely under defect as under no defect. Multiplying gives posterior odds four to nine. Converting those odds to probability gives four out of thirteen, again about 30.8 percent. Writing four out of nine would use the alternative count alone as the probability denominator and overstate the result.

The flag supplies evidence for defect because it raises the model probability from ten percent. It does not establish defect as certain or even more probable than absence. Further observations could change the assessment, but a duplicate copy of the same flag does not add independent evidence. An audit should retain the model assumptions alongside the calculation rather than presenting the percentage as a fact about every future batch.

10. Where this goes wrong

Reading the accuracy as the answer. The test is 98 percent right, so my positive is 98 percent reliable. Those are two different fractions with two different groups on the bottom.

Ignoring the size of the healthy row. A 2 percent slip-up rate sounds tiny until it is 2 percent of 9,990 people, which is 199.8 expected false positives, or about 200 — roughly twenty times the whole top row.

Concluding that the test is useless. It is not. Before the test your chance was 1 in 1,000; afterwards it is 1 in 20 in the illustrated table. The conditional probability is fifty times the base rate, and it is still far below certainty.

Thinking a bigger study fixes it. Test a million people instead and every number multiplies by a hundred. The share stays at 5 percent, because the share never depended on how many were tested.

11. A rare defect among flagged components

  1. Name the full batch and specified defect.

    10,000 components; 5 defective.

    The base rate concerns this batch and this property.

  2. Record both kinds of positive result.

    5 true flags and 245 false flags.

    Positive results arise from both underlying states.

  3. Count all flags.

    5 + 245 = 250.

    Every positive result belongs in the conditioning group.

  4. Calculate the defective share among flags.

    5/250 × 100 = 2%.

    The numerator is true positives, not all inspected components.

  5. Compare with the prior share.

    5/10,000 = 0.05%; the flagged share is 2%.

    The flag is informative in this record while remaining far from certainty.

12. Translate probability to odds and back

  1. Start with the prior probability.

    P(defect) = 1/10.

    The alternative has probability 9/10.

  2. Convert to prior odds.

    (1/10)/(9/10) = 1/9, or 1:9.

    Odds compare the two alternatives rather than one with the whole.

  3. Calculate the positive-result likelihood ratio.

    0.8/0.2 = 4.

    The flag is four times as probable under defect.

  4. Multiply odds by evidence strength.

    Posterior odds = (1/9) × 4 = 4/9, or 4:9.

    The update preserves the role of the initial base rate.

  5. Convert the new odds into probability.

    4/(4 + 9) = 4/13 ≈ 30.8%.

    Probability uses the combined weight of both alternatives.

13. Verify an odds update with a complete frequency table

  1. Represent the prior using a modeled batch.

    1,000 components: 100 defective, 900 not.

    The two groups instantiate the ten-percent prior.

  2. Apply the sensitivity to defects.

    80 flags and 20 misses.

    The positive and negative results must sum to one hundred.

  3. Apply the false-positive rate to nondefects.

    180 flags and 720 negatives.

    The same result is computed separately within the alternative group.

  4. Count the positive group.

    80 + 180 = 260 flags.

    The posterior conditions on all flags regardless of underlying state.

  5. Calculate the conditional share.

    80/260 = 4/13 ≈ 30.8%.

    This agrees with the odds update because both use identical inputs.

  6. Check the negative group as well.

    20/(20 + 720) = 1/37 ≈ 2.7% defective among negatives.

    A negative result is informative without logically excluding defect.

14. A leak detector's positive group

  1. Record the true and false flags.

    50 leaking sites flagged; 450 nonleaking sites flagged.

    The scenario supplies independent outcome classifications.

  2. Count all flagged sites.

    50 + 450 = 500.

    The denominator must include false alarms.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Compute the true share and interpret its limit.

15. Guided practice

A screening program checks ten thousand people at a free clinic for an eye condition that is caught early, using a photograph of the back of the eye. $40$ of them have it, and the test finds every one. Of the $9960$ who do not have it, $360$ still come out positive. Fill in the table.

Tests positiveTests negativeTotal
Has an eye condition that is caught early40
Does not have it9960
Total10000

16. Guided practice

Prior odds for a fictional fault are 1:9. A flag has probability 0.6 with the fault and 0.2 without it. Complete the odds update.

  1. Compute the evidence likelihood ratio.

    0.6 divided by 0.2 gives ratio.

    The ratio compares the same flag under the two alternatives.

  2. Apply the ratio to prior odds.

    Posterior odds with denominator nine are numerator:9.

    Multiplication changes the relative weight assigned to the fault.

  3. Convert the resulting odds to a percentage.

    Fault weight divided by total weight gives percent percent.

    Probability uses the combined weight of fault and no fault.

17. Guided practice

Among ten thousand schoolchildren, $10$ have a rare pollen allergy and all of them test positive. A further $190$ people who do not have it also test positive. For each person who really has it, how many people in total are told they have tested positive?

Answer:

18. Practice

Among ten thousand houses on one network, $50$ have a leak in the pipe under a house and test positive, and $450$ people who do not have it also test positive. Someone has just tested positive. Out of every hundred people in that position, how many really have it? Give a whole number of percent.

Answer:

19. Practice

A fictional defect model assigns prior odds 1:4 for defect versus no defect. A flag has probability 0.75 under defect and 0.25 under no defect. Give the likelihood ratio, the numerator of posterior odds written with denominator four, and the posterior probability as a fraction.

Likelihood ratio ratio; posterior odds numerator:4; posterior probability probability.

20. Somewhere new

A shop's tag alarm is checked over $10000$ shoppers. $15$ of them are shoplifting, and the alarm catches every one. It also sounds for $150$ honest shoppers. For each shoplifter the alarm catches, how many honest shoppers does it stop?

Answer:

21. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

22. Test question

A fictional model has 20% defective components. A flag occurs with probability 0.9 under defect and 0.1 under no defect. In a modeled batch of 1,000, give the expected true flags, expected false flags, and the posterior defect probability as a fraction.

True flags true_flags; false flags false_flags; posterior probability posterior.

23. What you can do now

You can build the screening table from a base rate and an error rate, and say what share of positive results are real. Tell someone why a test that catches every real case can still be wrong nineteen times out of twenty when it says yes. Next: what an accuracy figure on its own actually claims.

Working for the steps left to you

14. A leak detector's positive group, step 3

50/500 = 10% have the specified leak.

This conditional proportion does not determine whether any particular response is worthwhile.