Back to the on-screen lesson ·
Prior families the data leave in place, so a posterior is arithmetic on parameters, and the weight a prior still carries as the sample grows.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will update a beta prior with a count of successes and failures, read the posterior mean as a weighted average of the prior mean and the data, recognise which prior and likelihood pairs are conjugate, and say how many observations a prior is worth.
The last lesson's updates were over two or three hypotheses, so normalising meant adding a few numbers. Over a continuous parameter it means an integral, and for most prior and likelihood pairs that integral has no closed form. This lesson is about the pairs where it does.
A prior family is conjugate to a likelihood when the posterior lies in the same family, with different parameters. The beta family lives on the interval from zero to one and is conjugate to counts of successes and failures; the gamma family lives on the positive numbers and is conjugate to Poisson counts. A prior's parameters act as a prior sample size: how many observations it is worth.
A prior family is conjugate to a likelihood when prior times likelihood has the prior's own shape. The posterior is then read off by moving parameters, and no integral is ever computed.
Beta and binomial. A Beta$(a, b)$ prior on a success probability, meeting $s$ successes and $f$ failures, gives
$$\text{Beta}(a + s,\; b + f).$$
The successes go to the first parameter and the failures to the second. The posterior mean is $\dfrac{a + s}{a + b + s + f}$, which is a weighted average of the prior mean $\dfrac{a}{a+b}$ and the observed proportion $\dfrac{s}{s+f}$, weighted by $a + b$ against $s + f$. So $a + b$ is the prior sample size: the number of observations the prior is worth.
Normal and normal. A normal prior on a mean, with normal data of known spread, gives a normal posterior whose mean is the precision-weighted average of the prior mean and the sample mean, and whose precision is the sum of the two precisions. Information adds.
Gamma and Poisson. A gamma prior on a rate, meeting counts over a known length of time, gives a gamma posterior with the counts added to the shape and the time to the rate parameter.
Conjugacy is a property of the pair. A normal prior on a success probability is conjugate to nothing: it puts mass outside $[0, 1]$, and the product with a binomial likelihood has no standard form. That is not a broken calculation — the posterior exists and is perfectly well defined — but computing with it needs a numerical method rather than arithmetic. Modern practice does exactly that; conjugate families remain the way the idea is understood, and the way a posterior is checked by hand.
Another way: story
You are handed a coin and told nothing. Beta$(2, 2)$ says: assume somebody already tossed it four times and got two heads. Then every real toss is simply added to that imaginary record. After four hundred real tosses the imaginary four are one per cent of the evidence and nobody can tell they were ever there.
Another way: steps
| Prior | Data | Posterior | What the data add |
|---|---|---|---|
| Beta($a$, $b$) | $s$ successes, $f$ failures | Beta($a+s$, $b+f$) | counts, one to each parameter |
| Normal($m$, precision $r$) | normal, known spread | Normal, precision $r + nr_0$ | precision, which adds |
| Gamma(shape $a$, rate $c$) | Poisson counts over time $t$ | Gamma($a + \text{count}$, $c + t$) | the count and the time |
Every row is addition, and in every row the prior's parameters are in the same currency as the data: counts, precision, or time observed. That is what makes a conjugate prior interpretable rather than merely convenient — you can always say how many observations it is worth.
Crossing the two counts over. Successes go to the first parameter and failures to the second. Both are additions of small numbers, which is exactly why the slip is easy.
Reading a posterior parameter as a probability. Beta$(9, 5)$ has mean $9/14$, not $9$. The parameters are counts.
Believing conjugacy is a property of the prior. It is a property of the prior and the likelihood together.
Treating a non-conjugate case as an error. The posterior exists; it merely has no name. That is a computational inconvenience and nothing more.
Choosing a conjugate prior because it is convenient, and not saying so. Convenience is a reason to prefer one prior among several defensible ones. It is not a reason to believe it, and the choice belongs in the report.
Prior Beta($2$, $2$); $7$ successes in $10$ trials, so $3$ failures.
Prior worth four observations.
Posterior Beta($2 + 7$, $2 + 3$) = Beta($9$, $5$).
Counts added.
Posterior mean $9/14 \approx 0.64$, between the prior mean $0.5$ and the observed $0.7$.
A weighted average.
Prior $N(50, 10^{2})$; one observation of $70$ with a known spread of $10$.
Equal precisions.
The posterior mean is the equally weighted average, $60$.
Precision-weighted.
The posterior variance is $\dfrac{1}{1/100 + 1/100} = 50$, half the prior's: information added.
Precisions add.
The successes go to the first parameter: $3 + 5$.
First parameter.
The failures go to the second: $1 + 4$.
Second parameter.
So the posterior is Beta($8$, $5$), with mean $8/13$.
A beta prior on a success probability has parameters $5$ and $3$. Then $6$ successes and $2$ failures are observed. Give the two parameters of the posterior, and the total of the two.
| Value | |
|---|---|
| First parameter of the posterior | |
| Second parameter of the posterior | |
| The two added together |
A beta prior on a success probability has parameters $1$ and $1$, and then $7$ successes and $2$ failures are observed. The posterior is a beta with parameters $A$ and $B$. What is $A$?
Answer:
A beta prior with parameters $4$ and $4$ meets $4$ successes and $6$ failures. Give the two parameters of the posterior.
First parameter: p. Second parameter: q.
A prior with parameter $5$ is to be updated. Match each conjugate pair to the posterior it produces.
| A beta posterior, with the counts added to the two parameters | A normal posterior, its mean a precision-weighted average | A gamma posterior, with the counts and the time added to the two parameters | A posterior in no standard family, needing a numerical method | |
|---|---|---|---|---|
| A beta prior on a success probability, with counts of successes and failures | ||||
| A normal prior on a mean, with normal data of known spread | ||||
| A gamma prior on a rate, with counts over a known length of time | ||||
| A normal prior on a success probability, with counts of successes and failures |
A beta prior with both parameters equal to $2$ has mean $0.5$. A trial is run with the proportion of successes fixed at $0.75$. Give the posterior mean after $16$, $36$ and $96$ observations.
| Posterior mean | |
|---|---|
| After $16$ observations | |
| After $36$ observations | |
| After $96$ observations |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Consider a uniform prior on a normal mean, updated by normal observations with an unknown spread. Does the posterior stay in the prior's own family?
You can move the parameters of a conjugate prior to those of the posterior, and say what weight the prior still carries after a given amount of data. Say in your own words why conjugacy is a property of the prior and the likelihood together.
9. Your turn: prior Beta($3$, $1$), then $5$ successes and $4$ failures, step 3