Back to the on-screen lesson ·

Credible intervals

A region of the posterior holding a stated probability, the sentence it licenses about the parameter, the prior it costs, and the guarantee it does not offer.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will build a credible interval from a posterior, state the probability claim it supports about the parameter, and set that claim beside what a confidence interval does and does not assert. You will also say what the credible interval costs, what it gains, and which of the two a given purpose calls for.

2. The statement unit two could not make

Unit two built intervals and then spent a lesson saying carefully what they did not claim. The reason was that the parameter was not random, so no probability could be attached to it. This unit made it random. The statement is now available, and the last thing to do is to see exactly what it cost.

3. Words for this lesson

A credible interval at level $1 - \alpha$ is a region holding that much posterior probability. A central interval leaves $\alpha/2$ in each tail; a highest-density interval instead takes the most probable values, and is the shorter of the two whenever the posterior is skewed. Coverage is the frequentist property a confidence interval has and a credible interval does not promise.

4. A region of the posterior, and what that buys

Once the posterior $\pi(\theta \mid x)$ exists, an interval for $\theta$ is a region of it. A central $95\%$ credible interval runs from the $2.5$th to the $97.5$th percentile of the posterior, and it supports the sentence

given the prior and the data, there is a $95\%$ probability that $\theta$ lies in this interval.

That is the sentence people have always wanted to say about a confidence interval and never could. A confidence level is a property of the procedure, fixed before the data arrive. Once the numbers are in, the interval is a pair of numbers and the parameter is a number: it either covers or it does not, and no probability is left to report. The statement the level supports is about a long run of samples built the same way.

The price is the prior. Two analysts with different priors report different credible intervals from identical data, and neither is making an error. The prior has to be stated, and a report that hides it is hiding an input.

In practice the two intervals often nearly coincide. With a weak prior and plenty of data the posterior is dominated by the likelihood, and the credible interval and the confidence interval land in almost the same place. Coinciding numbers do not make the two statements interchangeable — what each claims is decided by how it was built, not by where it fell.

The credible interval has two practical advantages worth naming. It respects the parameter's range: a beta posterior on a proportion produces an interval inside $[0, 1]$ automatically, where a normal approximation can run below zero. And it composes: the posterior from one study is the prior for the next, so evidence accumulates without any new machinery.

It has one real disadvantage. It offers no guarantee about how often it covers over repeated samples, because that is not a question it asks. Where such a guarantee is the point — a regulator setting a standard, a process being monitored — the frequentist interval is the one that answers.

Another way: picture

The posterior drawn as a curve over the parameter, with the middle $95\%$ of its area shaded and a tail of $2.5\%$ left white at each end. The interval is the width of the shading. Nothing is repeated and no other sample is imagined; the probability is read straight off the area, which is what a confidence interval has no curve to do.

Another way: steps

  1. Obtain the posterior.
  2. Choose the level, and whether the interval is central or highest-density.
  3. Read off the percentiles, or use the posterior mean and standard deviation when it is near normal.
  4. Report it together with the prior it came from.

5. The two intervals, side by side

Confidence intervalCredible interval
The parameter isfixedrandom
What is randomthe intervalthe parameter
Needs a priornoyes
Probability about $\theta$not availableavailable
Guarantee over repeated samplesyesnot claimed
Respects the parameter's rangenot alwaysautomatically
Two analysts agreeyesonly if their priors agree

Every row is a genuine trade rather than a defect on one side. The last two rows are the ones that decide in practice: a regulator wants the guarantee and will accept the awkwardness, while an analyst combining several studies wants the composability and will state a prior to get it.

6. Where this goes wrong

Reading a confidence interval as a credible one. This is the single commonest misstatement in applied statistics. A confidence level is a property of the procedure, fixed before the data arrive. Once the numbers are in, the interval is a pair of numbers and the parameter is a number: it either covers or it does not, and no probability is left to report. The statement the level supports is about a long run of samples built the same way.

Believing a credible interval has a coverage guarantee. It does not claim one. Where the guarantee is what matters, use the interval that offers it.

Reporting a credible interval without the prior. The prior is an input. Omitting it makes the interval unreproducible.

Treating a flat prior as no prior. It is a prior, it is not flat after a change of variable, and on an unbounded parameter it may not even be a distribution.

Concluding from coinciding numbers that the two are the same. What an interval claims is decided by how it was built.

7. A credible interval from a beta posterior

  1. Prior Beta($2$, $2$), then $7$ successes in $10$ trials: posterior Beta($9$, $5$).

    Conjugate updating.

  2. Its mean is $9/14 \approx 0.64$, and its central $95\%$ region runs from about $0.39$ to $0.86$.

    Percentiles of the posterior.

  3. Reported as: given this prior and these data, there is a $95\%$ probability the success rate lies between $0.39$ and $0.86$.

    Both ends inside zero and one.

8. The same data, the frequentist interval

  1. $\hat p = 0.7$ from $10$ trials; standard error $\sqrt{0.7 \times 0.3/10} \approx 0.145$.

    No prior anywhere.

  2. A $95\%$ interval of $0.7 \pm 0.284$: from $0.416$ to $0.984$.

    Close to the credible one, and not the same.

  3. It claims coverage over repeated samples and says nothing about the probability that this one covers.

    Two intervals, two different claims.

9. Your turn: a posterior that is approximately normal with mean $40$ and standard deviation $2$, at the $95\%$ level

  1. The half-width is $1.96 \times 2$.

    Multiplier times posterior spread.

  2. That is $3.92$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    So the credible interval is $(36.08, 43.92)$, and it may be reported as a probability about the parameter.

10. Guided practice

After the data, the posterior for a parameter is approximately normal with mean $21$ and standard deviation $5$. Using a multiplier of $1.960$, give the central $95\%$ credible interval, with both endpoints included.

This task has no paper form; do it on a device.

11. Guided practice

A $95\%$ confidence interval has been computed from a sample. Is this claim about it correct: there is a ninety-five per cent probability that this particular interval contains the parameter?

12. Practice

A credible interval is built to hold $90\%$ of the posterior. Give the posterior probability inside the interval and the posterior probability outside it, both as proportions.

Inside: a. Outside: b.

13. Practice

An analysis of a parameter whose posterior mean is $48$ produces the four objects below. Match each to what it is.

A distribution over the parameter, before the dataA distribution over the parameter, after the dataA region holding a stated probability of one of those distributionsThe output of a procedure with a stated coverage over repeated samples
The prior
The posterior
A credible interval
A confidence interval

14. Somewhere new

A report gives both a $95\%$ confidence interval and a $95\%$ credible interval for a parameter, and both run from $33$ to $43$. Mark the two sentences that are true only of the credible interval.

This task has no paper form; do it on a device.

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A posterior for a success probability is a beta distribution whose parameters are $15$ and $5$. What is its mean, which a credible interval would be centred near?

Answer:

17. What you can do now

You can build a credible interval, say exactly what it claims, and tell that claim from a confidence interval's. Say in your own words what a credible interval buys and what it pays for it.

Working for the steps left to you

9. Your turn: a posterior that is approximately normal with mean $40$ and standard deviation $2$, at the $95\%$ level, step 3