Back to the on-screen lesson ·

The central limit theorem

Why a sum or an average of many independent contributions is nearly normal, what its mean and spread are, and the three hypotheses the statement needs.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will find the mean and standard deviation of a sum and of an average of independent measurements, standardise an observed average with the aggregate's own spread, and state the normal approximation the theorem supplies. You will also decide which claims the theorem reaches: aggregates rather than single values, approximately rather than exactly, and only with a finite variance.

2. What the law of large numbers left out

The law said the average settles at $\mu$ and that its variance is $\sigma^{2}/n$. It said nothing about the shape of the average's distribution — only how narrow it gets. That shape is what this lesson supplies, and it is the same shape for every starting distribution.

3. Words for this lesson

Standardising subtracts an aggregate's own mean and divides by its own standard deviation. Convergence in distribution means the cumulative distribution functions converge, which is a statement about shape rather than about what any particular run does. The standard error is the standard deviation of an average, $\sigma / \sqrt{n}$.

4. The shape a sum takes, whatever it was made of

Let $X_1, X_2, \ldots$ be independent with the same distribution, mean $\mu$ and finite variance $\sigma^{2}$. Then

$$\frac{\bar{X}_n - \mu}{\sigma / \sqrt{n}} \longrightarrow N(0, 1)$$ in distribution as $n \to \infty$, and equivalently the sum $\sum_{i \le n} X_i$ is approximately $N(n\mu, n\sigma^{2})$.

Read what that does and does not say. The average and the sum become nearly normal; a single measurement does not change at all. The starting shape is forgotten entirely except through its mean and its variance — which is why one table of normal probabilities serves errors of measurement, totals of insurance claims and averages of survey answers alike.

Three hypotheses earn their place. Independence is what makes the variances add. A finite variance is essential: the Cauchy distribution has none, and the average of $n$ Cauchy variables has exactly the same distribution as one of them, however large $n$ is. And the conclusion is an approximation, not an identity: how large $n$ must be before it is any good depends on how lopsided the starting distribution is.

The scaling is worth reading carefully. The sum's mean grows like $n$ while its standard deviation grows only like $\sqrt{n}$, so a total becomes relatively more predictable as it grows even as its absolute spread widens. That single sentence is why insurance works.

Another way: picture

The lopsided mass function of one die-and-coin contraption, then the distribution of the average of two of them, then of five, then of thirty. The first is spiky and one-sided; by thirty it is a symmetric bell, narrower than any of them, sitting over the same mean.

Another way: steps

  1. Check independence and a finite variance.
  2. Write the mean and the standard deviation of the aggregate: $n\mu$ and $\sqrt{n}\sigma$ for a sum, $\mu$ and $\sigma/\sqrt{n}$ for an average.
  3. Standardise with those numbers, not with one measurement's.
  4. Read the standard normal table, and say that the answer is an approximation.

5. Sum and average, side by side

One measurement has mean $\mu$ and standard deviation $\sigma$. For $n$ of them:

QuantityMeanStandard deviationAs $n$ grows
One measurement$\mu$$\sigma$unchanged
The sum$n\mu$$\sqrt{n}\,\sigma$spreads, but slower than it grows
The average$\mu$$\sigma/\sqrt{n}$tightens

Both aggregates are nearly normal and the single measurement is not. The middle row is the insurer's row: a book of a million policies has a total claim whose spread is a thousand times one policy's, against premiums a million times one policy's — so the relative uncertainty falls by a factor of a thousand.

6. Where this goes wrong

Applying it to a single measurement. Collecting a thousand lopsided measurements leaves each of them exactly as lopsided as before. Only their sum and average change shape.

Standardising with the wrong spread. An average of $n$ has standard deviation $\sigma/\sqrt{n}$, not $\sigma$. Using the wrong one makes every sample look ordinary.

Reading it as exact. It is a limit, and at any finite $n$ it is an approximation whose quality depends on the starting shape.

Forgetting the finite variance. Without it the theorem is simply false, and the Cauchy distribution is the standing counterexample: its averages never settle and never become normal.

7. A lopsided measurement, averaged

  1. Service times have mean $4$ minutes and standard deviation $4$, with a long right tail.

    One service time is not normal at all.

  2. The average of $100$ has mean $4$ and standard deviation $4/\sqrt{100} = 0.4$.

    Mean unchanged, spread divided by ten.

  3. So an average above $4.8$ is two standard deviations out: roughly a $2.5\%$ chance, from a normal table.

    The table applies to the average, not to one service.

8. Why a total becomes relatively predictable

  1. One policy pays out $200$ on average with standard deviation $1000$.

    Wildly uncertain on its own.

  2. A million policies: mean $2 \times 10^{8}$, standard deviation $1000\sqrt{10^{6}} = 10^{6}$.

    $\sqrt{n}$ against $n$.

  3. The total is within about half a per cent of its mean, from policies that were individually unpredictable.

    This is the whole business model.

9. Your turn: measurements have mean $20$ and standard deviation $6$; find the mean and standard deviation of the average of $36$, and standardise an average of $21$

  1. The average has mean $20$ and standard deviation $6/\sqrt{36} = 1$.

    Divide by the square root of the sample size.

  2. $z = (21 - 20)/1$.

    Subtract, then divide by the average's own spread.

  3. Your turn: work this step out. Its working is at the end of the packet.

    So $z = 1$: one standard deviation above, which is unremarkable.

10. Guided practice

One hundred independent measurements each have mean $80$ and standard deviation $20$. Give the mean and standard deviation of their total, and of their average.

Value
Mean of the total
Standard deviation of the total
Mean of the average
Standard deviation of the average

11. Guided practice

One hundred independent measurements each have mean $30$ and standard deviation $30$. Their average came out as $23$. How many standard deviations of the average is that from its mean?

Answer:

12. Practice

One hundred independent measurements each have mean $60$ and standard deviation $10$, with an unknown and lopsided distribution. Give the mean and standard deviation of the normal curve that approximates the distribution of their average.

The distribution of the average of a hundred measurements

Mean of the normal:

Standard deviation of the normal:

13. Practice

Measurements have mean $80$ and standard deviation $10$. Give the mean and standard deviation of the average of four of them.

Mean: p. Standard deviation: q.

14. Somewhere new

Sort these four statements about $789$ independent measurements by whether the central limit theorem supports them.

Supported by the theoremNot supported by the theorem
The total of the $789$ is approximately normal
The average of the $789$ is approximately normal
A single measurement is approximately normal
The average of the $789$ is exactly normal

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A machine's individual output weights are strongly lopsided, with a long tail to the right. What does the central limit theorem say about $482$ of them?

17. What you can do now

You can compute the mean and spread of a sum and of an average, standardise an average correctly, and name what the theorem does not claim. Say in your own words why a total becomes relatively more predictable as it grows even though its absolute spread widens.

Working for the steps left to you

9. Your turn: measurements have mean $20$ and standard deviation $6$; find the mean and standard deviation of the average of $36$, and standardise an average of $21$, step 3