Back to the on-screen lesson ·

An interval for a proportion

A proportion as a mean of zeros and ones, a standard error that comes free with the estimate, and the worst case a survey is planned on.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will estimate a population proportion from a survey, compute its standard error from the estimate itself, and build a confidence interval around it. You will also find the worst case a survey of a given size can face, and use it to plan a sample size before any data exist.

2. A mean in disguise

A proportion is the mean of a sample of zeros and ones, so everything in the last two lessons applies. What is new is that the spread of such a population is not a separate unknown: it is determined by the proportion itself, which changes what a survey can plan for before it has any data.

3. Words for this lesson

The sample proportion is the count of successes over the number of trials. Its standard error is $\sqrt{\hat p(1 - \hat p)/n}$, computed from the estimate itself. The worst case is the value of $p$ maximising that, which is a half, and a survey planned on it is planned for whatever answer arrives.

4. The spread comes free with the estimate

A single yes-or-no response has mean $p$ and variance $p(1 - p)$. The sample proportion $\hat p$ is the mean of $n$ such responses, so

$$E[\hat p] = p, \qquad \operatorname{Var}(\hat p) = \frac{p(1 - p)}{n},$$

and for large $n$ the central limit theorem makes $\hat p$ approximately normal. The interval is

$$\hat p \pm z_{\alpha/2}\sqrt{\frac{\hat p(1 - \hat p)}{n}}.$$

The variance did not have to be estimated separately: it is a function of the parameter, so plugging in $\hat p$ supplies it. That is why a proportion needs no t distribution in practice, and why a survey of a thousand people can quote a margin of error without knowing anything about the population's spread.

The function $p(1 - p)$ is a downward parabola, zero at both ends and largest at $p = \tfrac12$ where it is $\tfrac14$. So the worst case standard error is $\tfrac{1}{2\sqrt n}$, and a survey planned on that assumption is planned for every possible answer. At $n = 1000$ the worst-case margin at $95\%$ is about $3$ per cent — the number quoted beside almost every published poll.

The approximation needs enough data in both categories. The usual check is at least ten successes and ten failures; below that the interval can run outside $[0, 1]$, which is a sign to use an exact method instead.

Another way: story

A pollster has to book a sample size before a single person is asked, and the precision they can promise depends on an answer they do not yet have. The parabola rescues them: whatever the answer, the standard error is at most $1/(2\sqrt n)$. So they promise the worst case, and deliver at least it.

Another way: steps

  1. Compute the sample proportion.
  2. Compute its standard error from that proportion and the sample size.
  3. Multiply by the critical value for the level.
  4. Add and subtract, and check both ends lie inside $[0, 1]$.

5. The worst case, and what it costs to beat it

Worst-case $95\%$ margins of error, at a critical value of $1.96$.

Sample sizeWorst-case standard errorMargin of error
1000.059.8%
4000.0254.9%
10000.01583.1%
25000.012.0%

The third row is the industry standard, and the table says why: it is the smallest sample that gets the margin under about three points. Halving it again would take four thousand people, and almost no question is worth that.

6. Where this goes wrong

Thinking the standard error is largest at the extremes. It is smallest there and largest at a half, because $p(1 - p)$ is a downward parabola.

Reading the worst case as this survey's standard error. The worst case is a planning figure that ignores the data; the survey's own standard error uses $\hat p$ and is usually smaller.

Quoting a margin of error as though it applied to every question on the survey. It does, as a worst case, and it overstates the uncertainty badly for any question answered by five per cent of people.

Using the normal interval with very few successes. With three successes in two hundred the interval can run below zero, which is a sign the approximation has failed rather than a value to report.

7. A survey, reported

  1. $620$ of $1000$ people say yes: $\hat p = 0.62$.

    Count over count.

  2. Standard error $\sqrt{0.62 \times 0.38 / 1000} \approx 0.0154$; at $95\%$ the margin is about $0.030$.

    From the estimate itself.

  3. Interval about $(0.590, 0.650)$, or $59\%$ to $65\%$.

    Both ends inside zero and one.

8. Planning before the data

  1. A margin of $2$ percentage points is wanted at $95\%$, and nothing is known about the answer.

    So assume the worst case.

  2. $1.96 \times 0.5/\sqrt n = 0.02$ gives $\sqrt n = 49$, so $n = 2401$.

    Solve for the sample size.

  3. Round up to $2500$ and the promise holds whatever the answer turns out to be.

    Rounding up is the safe direction.

9. Your turn: $90$ of $400$ say yes, at a critical value of $2$

  1. $\hat p = 90/400 = 0.225$.

    Count over count.

  2. The standard error is $\sqrt{0.225 \times 0.775 / 400} \approx 0.0209$, so the margin is about $0.0418$.

    From the estimate itself.

  3. Your turn: work this step out. Its working is at the end of the packet.

    The interval is about $(0.183, 0.267)$.

10. Guided practice

In a survey of $200$ people, $69$ answered yes. What is the estimate of the population proportion answering yes?

Answer:

11. Guided practice

A survey of $200$ people found $77$ saying yes, and the standard error of that proportion is $1/25$. Using a critical value of $2$, give the confidence interval for the population proportion, with both endpoints included.

This task has no paper form; do it on a device.

12. Practice

A survey of $2500$ people is planned. For each candidate value of the population proportion, give the product of that proportion with one minus it, and the resulting standard error.

Proportion times one minus itStandard error
If the proportion is $0.5$
If the proportion is $0.2$
If the proportion is $0.1$

13. Practice

A survey of $400$ people found $246$ saying yes. Give the sample proportion, and the largest the standard error could have been for a survey of this size, whatever the answer had turned out to be.

Sample proportion: p. Worst-case standard error: c.

14. Somewhere new

A survey of $400$ people found $98$ saying yes, giving an interval of $39/200$ to $59/200$. A second survey of $1600$ people found the same proportion. Give the second survey's interval, with both endpoints included.

This task has no paper form; do it on a device.

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A survey of $2500$ people estimates a proportion. At which value of that proportion is the standard error of the estimate largest?

17. What you can do now

You can build an interval for a proportion, say where its standard error is largest, and plan a survey on the worst case. Say in your own words why a proportion needs no separate estimate of the population spread.

Working for the steps left to you

9. Your turn: $90$ of $400$ say yes, at a critical value of $2$, step 3