Back to the on-screen lesson ·
Turning a likelihood into a posterior with natural frequencies, and why the same test means different things in different populations.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will turn a likelihood into a posterior by counting natural frequencies in a chosen population, name the prior, the likelihood and the denominator in a Bayes calculation, and say what a positive result is worth. You will also show that the posterior depends on the prior by running one unchanged test across several populations.
The multiplication rule says $P(A \cap B)$ can be written two ways, and the law of total probability says the denominator of a conditional probability can be assembled from the branches. Bayes' theorem is those two facts written on one line; there is no new principle in it at all.
The prior is what you believed before the evidence; the likelihood is how well each hypothesis explains the evidence; the posterior is what you believe after. Sensitivity is $P(+ \mid \text{ill})$ and the false positive rate is $P(+ \mid \text{well})$. The base rate is the prior, under the name it usually goes by when it is being ignored.
For a partition $B_1, \ldots, B_n$,
$$P(B_j \mid A) = \frac{P(A \mid B_j) \, P(B_j)}{\sum_i P(A \mid B_i) \, P(B_i)}.$$
Every piece has a name. $P(B_j)$ is the prior — what you believed before the evidence. $P(A \mid B_j)$ is the likelihood — how well each hypothesis explains what you saw. The denominator is the total probability of the evidence, and it is only there to make the posteriors add to one. $P(B_j \mid A)$ is the posterior.
$P(A \mid B)$ and $P(B \mid A)$ are different numbers, and the difference is the whole of Bayes' theorem. Reading one for the other is the error behind almost every famous misuse of probability: a test that is right ninety-nine times in a hundred does not make a positive result ninety-nine per cent likely to be true.
The practical advice of this lesson is arithmetical rather than conceptual: work in natural frequencies. Take a thousand people, count how many are in each group, count how many of each group tests positive, and divide. The same theorem, and almost nobody gets it wrong that way.
Another way: story
A condition affects one person in a thousand. A test finds every case and is positive for one healthy person in twenty. Screen a hundred thousand people: a hundred have it and all test positive; of the remaining ninety-nine thousand nine hundred, about five thousand test positive too. So about one positive result in fifty is genuine — from a test that never misses a case.
Another way: steps
A test that finds $90\%$ of cases and is positive for $10\%$ of those without the condition, run on a thousand people each time.
| Prevalence | True positives | False positives | A positive is genuine |
|---|---|---|---|
| 1 in 100 | 9 | 99 | $1/12$ |
| 1 in 10 | 90 | 90 | $1/2$ |
| 1 in 2 | 450 | 50 | $9/10$ |
| 9 in 10 | 810 | 10 | $81/82$ |
Nothing in the middle two columns' rates changed across the rows. The last column runs from about $8\%$ to about $99\%$. A test has no accuracy independent of the population it is used on, which is why a screening programme and a diagnostic clinic report different things from the same instrument and neither is lying.
The base-rate fallacy. Reading the false positive rate, or one minus it, as the probability that a positive result is genuine. It ignores the prior entirely, and the prior is usually what dominates.
Reading sensitivity as the posterior. Finds every case is $P(+ \mid \text{ill}) = 1$. It says nothing at all about $P(\text{ill} \mid +)$.
Forgetting how many healthy people there are. A small percentage of a huge group beats a large percentage of a tiny one. That single sentence is most of the surprise in this lesson.
Multiplying percentages instead of counting people. The arithmetic is the same and the error rate is not. Count whole people.
Prevalence $1\%$, sensitivity $90\%$, false positive rate $10\%$, in a thousand people.
Choose the population first.
Ten have it, nine of them test positive. Nine hundred and ninety do not, and ninety-nine of those test positive.
Count each group separately.
So $9$ of $108$ positives are genuine: about $8\%$.
One division at the end.
$P(\text{ill}) = 0.01$, $P(+ \mid \text{ill}) = 0.9$, $P(+ \mid \text{well}) = 0.1$.
Prior and the two likelihoods.
$P(+) = 0.9(0.01) + 0.1(0.99) = 0.108$.
Total probability supplies the denominator.
$P(\text{ill} \mid +) = 0.009 / 0.108 = 1/12$.
The same answer, and harder to read.
Take ninety boxes: nine hold a prize and all nine beep.
Choose a population that divides cleanly.
The other eighty-one are empty, and one in nine of those beeps: nine boxes.
Count the false alarms.
So nine of the eighteen beeping boxes hold a prize: one half.
A thousand people are screened for a condition that one person in ten has. Of them, $100$ have the condition. The test finds it in $90\%$ of those who have it and gives a positive for $10\%$ of those who do not. Fill in the four figures.
| Value | |
|---|---|
| Have it and test positive | |
| Do not have it and test positive | |
| Test positive altogether | |
| Share of the positives who have it |
In a thousand people screened for a condition that one person in five has, screened with a test that never misses a case, $200$ have the condition. The test finds it in $100\%$ of those who have it and is positive for $25\%$ of those who do not. Someone tests positive. What is the probability they have the condition?
Answer:
A test detects a condition in every case that has it, and is positive for $9\%$ of people who do not. Someone tests positive. What is the probability that they have the condition?
A thousand people are screened for a condition that one person in five has, screened with a test that never misses a case; $200$ have it, the test finds $100\%$ of those cases and is positive for $25\%$ of the rest. How many positives of each kind are there?
True positives: p people. False positives: q people.
The same test throughout: it finds $90\%$ of cases and is positive for $10\%$ of those without the condition. Match each population to the probability that a positive result is genuine.
| $\dfrac{1}{12}$ | $\dfrac{1}{2}$ | $\dfrac{9}{10}$ | $\dfrac{81}{82}$ | |
|---|---|---|---|---|
| One person in a hundred has it | ||||
| One person in ten has it | ||||
| One person in two has it | ||||
| Nine people in ten have it |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
One test throughout: it finds $90\%$ of cases and is positive for $10\%$ of people who do not have the condition. Put these four populations in order by the probability that a positive result is genuine, smallest first.
Number the steps in order (write the number in the box):
You can compute a posterior by counting people in a round population, and name the prior and the likelihood in any such problem. Say in your own words why a test that never misses a case can still produce mostly false positives.
9. Your turn: one box in ten holds a prize; a detector beeps for every prize box and for one empty box in nine; find the chance a beeping box holds a prize, step 3