Back to the on-screen lesson ·

The law of large numbers

Why the variance of an average falls like one over the sample size, the weak law that follows from Chebyshev, and the fallacy the law is mistaken for.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will compute the mean and variance of an average of independent measurements, derive the weak law of large numbers from Chebyshev's inequality, and work out how many measurements a target precision costs. You will also say what the law does not claim: nothing about the next trial, and nothing about the counts catching up.

2. Two results already in hand

Variances of independent variables add, and Chebyshev bounds how far a variable strays in terms of its variance. Put them together and the law of large numbers falls out in three lines — which is why it appears here rather than as an article of faith.

3. Words for this lesson

The sample mean $\bar{X}_n$ is the average of the first $n$ measurements. Convergence in probability means the chance of being more than any fixed distance away goes to zero as $n$ grows. The weak law asserts exactly that; the strong law asserts that almost every single infinite run converges, which is a stronger claim about a different object.

4. The average is outvoted, not corrected

Let $X_1, X_2, \ldots$ be independent with the same distribution, mean $\mu$ and finite variance $\sigma^{2}$, and let $\bar{X}_n = \frac{1}{n}\sum_{i \le n} X_i$.

Mean and variance of the average. Linearity gives $E[\bar{X}_n] = \mu$ for every $n$. Independence makes the variances of the sum add, so $\operatorname{Var}(\sum X_i) = n\sigma^{2}$, and dividing by $n$ divides a variance by $n^{2}$:

$$\operatorname{Var}(\bar{X}_n) = \frac{\sigma^{2}}{n}.$$

The weak law. Chebyshev applied to $\bar{X}_n$ gives, for any $\varepsilon > 0$, $$P\!\left(|\bar{X}_n - \mu| \ge \varepsilon\right) \le \frac{\sigma^{2}}{n \varepsilon^{2}} \longrightarrow 0.$$ That is the whole proof: the average converges to $\mu$ in probability.

The strong law says more — the sequence of averages converges to $\mu$ with probability one, for every single run rather than in the aggregate — and needs only a finite mean. Its proof belongs to a later course; the distinction worth carrying is that the weak law bounds the chance of being off at each $n$, while the strong law describes what one infinite run does.

Every limit theorem in this course is a statement about a long-run average, and about nothing else. The coin has no memory, no debt and no plan; the average settles down because later trials outnumber the early ones, not because anything corrects for them.

Another way: story

A coin lands tails ten times. Toss it a million more times and the share of heads will be close to a half — not because the early tails are paid back, but because ten tosses out of a million and ten are almost nothing. The imbalance is still there; it has simply stopped mattering.

Another way: steps

  1. Check the trials are independent with the same distribution.
  2. Note that the average has mean $\mu$ at every $n$.
  3. Its variance is $\sigma^{2}/n$, which goes to zero.
  4. Chebyshev turns that into a probability, and the law follows.

5. What shrinks and what does not

A fair coin, with $H_n$ the number of heads in $n$ tosses. Typical sizes:

$n$Typical $\|H_n - n/2\|$Typical $\|H_n/n - 1/2\|$
10050.05
10 000500.005
1 000 0005000.0005

The middle column grows — it goes up like $\sqrt{n}$ — while the right column shrinks like $1/\sqrt{n}$. Both statements are true at once, and reading the first column as if it were the second is the gambler's fallacy in its most respectable dress. Nothing catches up; the denominator runs away.

6. Where this goes wrong

The gambler's fallacy. A run of tails does not make heads more likely. Independence is exactly the statement that it does not.

Expecting the counts to balance. The share converges; the difference between the counts typically grows without bound.

Reading it as a statement about one trial. It is about the average of many, and it says nothing about any particular one.

Forgetting the hypotheses. Independence, identical distribution, and a finite mean. A distribution with no mean — the Cauchy — has averages that never settle at all, however long the run.

7. The proof, in three lines

  1. $E[\bar{X}_n] = \mu$ by linearity, at every $n$.

    No independence needed for this line.

  2. $\operatorname{Var}(\bar{X}_n) = \sigma^{2}/n$ by independence.

    This line is where independence is spent.

  3. Chebyshev: $P(|\bar{X}_n - \mu| \ge \varepsilon) \le \sigma^{2}/(n\varepsilon^{2}) \to 0$.

    The weak law.

8. What it costs to halve the error

  1. With $\sigma = 10$, the average of $100$ measurements has standard deviation $1$.

    $\sigma/\sqrt{n}$.

  2. To halve that to $0.5$, $n$ must be $400$.

    Four times the work.

  3. Precision improves like $\sqrt{n}$, so the last decimal place is always the expensive one.

    The law's practical content.

9. Your turn: measurements have variance $36$; find the variance of the average of nine, and how many are needed for a variance of $1$

  1. $\operatorname{Var}(\bar{X}_9) = 36/9 = 4$.

    $\sigma^{2}/n$.

  2. For a variance of $1$, solve $36/n = 1$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    So $n = 36$ measurements.

10. Guided practice

Independent measurements each have variance $500$. Give the variance of their average for each sample size.

Variance
One measurement
Average of four
Average of twenty-five
Average of a hundred

11. Guided practice

Is this sound: the total of a thousand independent rolls is close to Normal in shape?

12. Practice

Independent measurements each have variance $500$. What is the variance of the average of $50$ of them?

Answer:

13. Practice

Measurements have variance $300$ each. Give the variance of the average of $50$ of them, and of the average of $200$.

With $50$: p. With $200$: q.

14. Somewhere new

Four sentences written after $12$ tosses of a fair coin all came up tails. Mark the two that are wrong.

This task has no paper form; do it on a device.

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

Measurements have variance $200$ each. Put these averages in order of how spread out they are, widest first.

Number the steps in order (write the number in the box):

17. What you can do now

You can show that the variance of an average is the single-measurement variance over the sample size, derive the weak law from Chebyshev, and reject a claim that makes the past correct itself. Say in your own words why the share of heads settles down while the difference between the counts does not.

Working for the steps left to you

9. Your turn: measurements have variance $36$; find the variance of the average of nine, and how many are needed for a variance of $1$, step 3