Back to the on-screen lesson ·

Priors, likelihoods and posteriors

Treating the parameter as random, multiplying a prior by a likelihood, and normalising the product into the distribution the data leave behind.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will carry out a Bayesian update over a small set of hypotheses, multiplying each prior by its likelihood and normalising the products into posterior probabilities. You will name the role each ingredient of Bayes' rule plays, say how a posterior differs from a likelihood, and run the update in the right order.

2. A different thing treated as random

Every unit so far has held the parameter fixed and let the data vary. That is what made a confidence level a statement about a procedure and a P-value a statement about data. This unit reverses the convention: the parameter is treated as random, and the price and the payoff of doing so are both large.

3. Words for this lesson

The prior is a distribution over the parameter before the data. The likelihood is the probability of the observed data as a function of the parameter — the same object as in unit one. The posterior is the distribution over the parameter after the data. The normalising constant is the total of prior times likelihood over all parameter values, and makes the posterior add to one.

4. Prior times likelihood, normalised

Treat the parameter $\theta$ as random with a prior $\pi(\theta)$. Bayes' rule gives the posterior

$$\pi(\theta \mid x) = \frac{L(\theta)\,\pi(\theta)}{\sum_{\theta'} L(\theta')\,\pi(\theta')} \;\propto\; L(\theta)\,\pi(\theta).$$

The denominator does not depend on $\theta$, so it is often left until last: recognise the shape of the numerator, then normalise.

The likelihood is the same object as in unit one, and this is where the difference between it and a posterior becomes sharp. $L(\theta)$ is a function of $\theta$ that does not integrate to one and assigns no probability to anything. Multiplying by a prior and normalising is precisely what turns it into a distribution — and the prior is the ingredient that makes that possible. There is no way to have the probability statement without it.

What is bought is considerable. A posterior supports the statement a confidence interval could not make: given the prior and the data, there is a $95\%$ probability that the parameter lies here. What is paid is that the prior must be stated, and two analysts with different priors reach different answers from identical data.

Two reassurances soften that. As data accumulate the likelihood sharpens and the prior's influence shrinks, so different priors converge. And the procedure is self-composing: today's posterior is tomorrow's prior, and a second batch of data runs the same four steps again, reaching the same place as analysing both batches at once.

Another way: picture

Three curves on one axis over the parameter. A broad prior, a narrow and off-centre likelihood, and between them the posterior — narrower than the prior and pulled towards the likelihood. With little data the posterior sits near the prior; with a great deal it sits almost on top of the likelihood, and the prior has become invisible.

Another way: steps

  1. State the prior, from something other than these data.
  2. Write the likelihood of the data that arrived.
  3. Multiply, value by value.
  4. Normalise so the result adds to one.

5. One update, in a table

Three candidate values, a prior, and data whose likelihood under each is shown.

ValuePriorLikelihoodProductPosterior
A0.50.10.050.25
B0.30.40.120.60
C0.20.150.030.15
1.000.201.00

The products add to $0.20$, which is the normalising constant, and dividing each by it gives the last column. $B$ began as the least likely of the three and ends as the most likely, because the data favoured it strongly enough to overcome the prior — which is what an update is for. Note also that the products are not probabilities: they total $0.20$, not $1$.

6. Where this goes wrong

Reading the likelihood as a probability for the parameter. It is a function of the parameter that does not integrate to one, and no probability can be read off it without a prior.

Forgetting to normalise. Products of prior and likelihood are not probabilities. Reporting one as though it were understates every posterior by the same factor.

Believing the prior drops out once there are data. Its influence shrinks as data accumulate and it never vanishes, which is why it must be stated.

Choosing a prior from the data. The prior is what was believed before. Fitting it to the same data it will be updated by counts the data twice.

Calling a flat prior 'no assumption'. It is an assumption, and one that is not even flat after a change of variable.

7. Two hypotheses, equal priors

  1. Two coins, one fair and one two-headed, chosen at random; three heads are observed.

    Equal priors of $0.5$.

  2. Likelihoods $1/8$ and $1$; products $1/16$ and $1/2$, totalling $9/16$.

    Prior times likelihood.

  3. Posteriors $1/9$ and $8/9$: the two-headed coin is now eight times as likely.

    Normalised.

8. The same data, a different prior

  1. Now only one coin in a hundred is two-headed; three heads again.

    Priors $0.99$ and $0.01$.

  2. Products $0.99/8 = 0.1238$ and $0.01$, totalling $0.1338$.

    The prior does real work.

  3. Posteriors $0.925$ and $0.075$: the fair coin is still far more likely, from identical data.

    Same data, different answer.

9. Your turn: equal priors, likelihoods $0.3$ and $0.1$

  1. The products are $0.15$ and $0.05$.

    Prior times likelihood.

  2. They total $0.2$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    So the posteriors are $0.75$ and $0.25$.

10. Guided practice

Two hypotheses are equally likely before the data. The data that arrived have likelihood $3$ under the first and $7$ under the second, in the same units. What is the posterior probability of the first?

Answer:

11. Guided practice

Two hypotheses have prior probabilities $0.25$ and $0.75$. The data that arrived have likelihood $5$ under the first and $5$ under the second, in the same units. Give each hypothesis's prior times its likelihood, and its posterior probability.

Prior times likelihoodPosterior probability
The first hypothesis
The second hypothesis

12. Practice

Two hypotheses are equally likely before the data, and the data have likelihood $9$ under the first and $1$ under the second, in the same units. Give the posterior probability of each.

First hypothesis: g. Second hypothesis: h.

13. Practice

A parameter is updated from data whose likelihood under one candidate value is $8$. Match each ingredient of the update to its role.

What was believed about the parameter before the data arrivedHow well each value of the parameter accounts for the data observedWhat is believed about the parameter after the dataThe total of the products, whose only job is to make the result add to one
The prior
The likelihood
The posterior
The normalising constant

14. Somewhere new

A parameter is to be updated by data whose likelihood under one candidate value is $1$. Put the steps of the update into the order they are carried out.

Number the steps in order (write the number in the box):

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

The likelihood of the observed data under one candidate value of a parameter is $1$. In what way does the posterior differ from the likelihood?

17. What you can do now

You can update a prior to a posterior over two or three hypotheses and name each ingredient of the rule. Say in your own words why a likelihood alone cannot give a probability for a parameter.

Working for the steps left to you

9. Your turn: equal priors, likelihoods $0.3$ and $0.1$, step 3