Back to the on-screen lesson ·

Confidence intervals

Point estimates and margins of error, the interval for a mean and a proportion, critical values from the normal and from Student's $t$, and what a confidence level does and does not claim.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you can build a confidence interval for a mean or a proportion, choose between a normal and a $t$ critical value on the right grounds, and say what widens an interval and what it costs to narrow one. You can also state what the confidence level means — a property of the method across repeated samples — and why it is not a probability about the interval you happen to have.

2. What you bring to this

You know that a sample mean has a sampling distribution centred on the population mean, with standard error $\sigma/\sqrt{n}$, and that it is approximately normal for a large enough sample. A confidence interval is that fact read backwards: instead of asking where the statistic will land, it asks which parameters could plausibly have produced the statistic that did land.

3. Words you will need

Point estimate: the single number computed from the sample — $\bar{x}$ or $\hat{p}$.

Margin of error: critical value $\times$ standard error.

Confidence interval: estimate $\pm$ margin.

Confidence level: the long-run proportion of such intervals that contain the parameter.

Critical value: $z^$ when $\sigma$ is known, $t^$ on $n-1$ degrees of freedom when it is estimated from the sample.

4. A range of values the data does not rule out

A point estimate is almost certainly wrong: the sample mean is not the population mean. What can be reported honestly is a range, together with how often the method that produced it succeeds.

Every confidence interval in this course has one shape:

$$\text{estimate} \pm \text{critical value} \times \text{standard error}$$

For a mean with $\sigma$ known, that is $\bar{x} \pm z^ \dfrac{\sigma}{\sqrt{n}}$. When $\sigma$ is unknown — which is nearly always — it is estimated by $s$, and the extra uncertainty is paid for by using Student's $t$ on $n - 1$ degrees of freedom instead of the normal: $\bar{x} \pm t^ \dfrac{s}{\sqrt{n}}$. For a proportion it is $\hat{p} \pm z^* \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}$.

The confidence level is where care is needed. It describes the method over repetition: build $95\%$ intervals from many samples and about $95\%$ of them will contain the parameter. It does not describe the one interval in front of you, because once computed there is nothing random left in it.

Three things widen an interval: a higher confidence level, a more variable population, and a smaller sample. Only the last is usually under anyone's control, and it costs four times the data to halve the width.

Another way: picture

Picture a hundred horizontal bars stacked up the page, each a $95\%$ interval from a different sample, and one vertical line through them all at the true mean. About ninety-five of the bars cross the line and about five miss it entirely. Nothing distinguishes a hit from a miss from the inside — which is precisely why the level describes the procedure and not any one bar.

Another way: steps

  1. Name the parameter and check the conditions — random sampling, and a large enough sample.
  2. Compute the point estimate.
  3. Compute the standard error.
  4. Look up the critical value: $z^$ if $\sigma$ is known, $t^$ on $n - 1$ degrees of freedom if it is estimated.
  5. Interval $=$ estimate $\pm$ critical value $\times$ standard error, and interpret it in context — about the parameter, not about the data.

5. What each choice costs

ChangeEffect on the interval
Level $90\% \to 99\%$wider (critical value $1.645 \to 2.576$)
Sample size $\times 4$half as wide
$\sigma$ unknown, estimatedwider, by using $t^$ instead of $z^$
More variable populationwider, in proportion to $\sigma$

Only one row is free of charge, and it is the one that says the interval must get wider.

6. Three things that trip people up

"There is a 95% chance the mean is in this interval." Once the interval is computed, both it and the parameter are fixed numbers: it either contains the mean or it does not. The $95\%$ is a property of the procedure, measured over the samples that might have been taken.

"95% of the data lies in the interval." The interval is about the population mean, and it is far narrower than the data. An interval of $[68, 72]$ for mean height does not claim that most people are between those heights.

"Higher confidence is better." Higher confidence is wider, and a wide enough interval says nothing. A $100\%$ interval covers every possible value and is useless. The level is a choice about which error to risk.

7. A $95\%$ interval for a mean, $\sigma$ known

  1. $\bar{x} = 50$, $\sigma = 12$, $n = 36$.

    Given.

  2. Standard error $= 12 / 6 = 2$, and $z^* = 1.96$ for $95\%$.

    The two factors.

  3. Margin $= 1.96 \times 2 = 3.92$, so the interval is $50 \pm 3.92$: $46.08$ to $53.92$.

    Estimate plus and minus.

  4. In context: we are $95\%$ confident the population mean lies between $46.08$ and $53.92$ — a statement about $\mu$, not about any person.

    Interpretation is part of the answer.

8. The same data with $\sigma$ unknown

  1. $s = 12$ from the sample, $n = 36$, so $df = 35$.

    Estimated spread.

  2. $t^* \approx 2.03$ rather than $1.96$ — a little larger.

    The price of estimating.

  3. Margin $= 2.03 \times 2 = 4.06$: a slightly wider interval, $45.94$ to $54.06$.

9. Your turn: $\bar{x} = 100$, $\sigma = 15$, $n = 25$, $90\%$ confidence

  1. Standard error $= 15 / 5 = 3$.

    Root of the sample size.

  2. $z^* = 1.645$ for $90\%$, so the margin is $1.645 \times 3 \approx 4.94$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    The interval is $95.06$ to $104.94$ — and at $99\%$ it would be wider, for the same data.

10. Guided practice

What is the normal critical value $z^*$ for a $98\%$ confidence interval?

answer

11. Guided practice

A sample of $49$ comes from a population with $\sigma = 42$. What is the margin of error for a $90\%$ confidence interval for the mean?

answer

12. Practice

A sample of $100$ has mean $76$, from a population with $\sigma = 30$. Give the lower and then the upper end of the $99\%$ confidence interval for the population mean.

from lo to hi

13. Practice

A study reports a $90\%$ confidence interval for a population mean. What does the $90\%$ describe?

14. Practice

A $98\%$ interval is recomputed at a higher confidence level, with everything else unchanged. What happens to it?

15. Practice

A sample of $6$ is taken and $\sigma$ is unknown, so the sample standard deviation is used instead. What critical value does a $95\%$ confidence interval need?

answer

16. Somewhere new

A manufacturer claims a mean of $91$. A study's $95\%$ confidence interval runs from $82$ to $88$. What can be said?

17. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

18. Test question

A report gives a confidence interval running from $72$ to $76$ but does not say what the sample mean was. What was it?

answer

19. What you can do now

You can build and interpret a confidence interval. Without looking: what are the two factors in a margin of error, and what exactly does the $95\%$ in a $95\%$ interval describe?

Working for the steps left to you

9. Your turn: $\bar{x} = 100$, $\sigma = 15$, $n = 25$, $90\%$ confidence, step 3