Back to the on-screen lesson ·
What an estimator is, why it is random, and how its spread falls with the size of the sample.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will tell a parameter from a statistic, name the estimator a question calls for, and compute the mean and the variance of the sample mean for any population with a finite variance. You will report an estimate beside its standard error, and say how much data a stated gain in precision costs.
Probability starts from a distribution and works out what the data will look like. Statistics runs the same machinery backwards: the data have arrived and the distribution is what is unknown. The one result carried over unchanged is that a sum of independent variables has the sum of their variances, which is where the whole of this lesson comes from.
A parameter is a fixed number describing the population — unknown, but not random. A statistic is any number computed from the sample. An estimator is a statistic used to target a parameter, and an estimate is the value it took on the sample that actually arrived. The sampling distribution is the distribution of an estimator over all the samples that might have been drawn, and its standard deviation is called the standard error.
Write $X_1, \dots, X_n$ for independent draws from a population with mean $\mu$ and variance $\sigma^{2}$. The sample mean
$$\bar X = \frac{1}{n}\sum_{i=1}^{n} X_i$$
is a function of the sample, so it is a random variable, and its distribution over repeated samples is its sampling distribution. Two facts about it hold for every population with a finite variance:
$$E[\bar X] = \mu, \qquad \operatorname{Var}(\bar X) = \frac{\sigma^{2}}{n}.$$
The first says the sample mean aims at the right place; the second says how tightly it clusters there. Its square root, $\sigma/\sqrt n$, is the standard error — the typical distance between an estimate and the parameter it estimates.
The square root is the whole economics of data collection. Four times the sample halves the standard error; a hundred times divides it by ten. Nothing about this depends on the population being normal: the mean and the variance of $\bar X$ are what they are for any population with a finite variance, and the central limit theorem then makes the shape approximately normal as well.
Another way: picture
Draw a sample, compute its mean, mark it on a line; rub out the sample and do it again, and again. The marks pile up into a hill centred on the population mean. Repeat the whole exercise with samples four times as large and the hill is half as wide, in exactly the same place.
Another way: steps
One population, standard deviation $20$, and the price of precision.
| Sample size | Standard error | Cost relative to the first row |
|---|---|---|
| 25 | 4 | 1 |
| 100 | 2 | 4 |
| 400 | 1 | 16 |
| 2500 | 0.4 | 100 |
Each halving of the standard error costs four times the data. Read the table the other way and it explains a fact about published surveys that otherwise looks like laziness: the margin quoted by a national poll of two thousand people would need forty thousand to be halved, and almost no question is worth that.
Treating the parameter as random because it is unknown. The population mean is a fixed number. Not knowing it is a fact about you, not about it.
Confusing the standard deviation with the standard error. The first describes how spread out the data are, the second how spread out the estimate is. They differ by a factor of $\sqrt n$, and the second shrinks while the first does not.
Expecting a larger sample to be closer to the truth every time. It is closer on average. Any particular large sample may land further out than any particular small one.
A population has standard deviation $12$; a sample of $36$ is drawn.
Name $\sigma$ and $n$.
$\operatorname{Var}(\bar X) = 144/36 = 4$, so the standard error is $2$.
Variance first, then its square root.
A sample mean of $51$ is reported as $51$ with a standard error of $2$.
Two numbers, never one.
The same sample's largest value is used to estimate the largest value in the population.
A statistic, since it is computed from the sample.
It can never exceed the population's largest value, so it lands below it on every sample.
Aiming low by construction.
Its sampling distribution sits entirely on one side of the target, which is a fault the mean does not have.
The next lesson is about exactly this.
The variance of the sample mean is $50/200$.
Divide by the sample size.
That is $0.25$.
So the standard error is $0.5$.
A sample of $4$ independent observations is drawn from a population with variance $12$. What is the variance of the sample mean?
Answer:
A population has standard deviation $2$. Give the standard error of the sample mean at each of these sample sizes.
| Standard error | |
|---|---|
| A sample of $4$ | |
| A sample of $16$ | |
| A sample of $100$ |
A sample of $81$ readings has mean $35$, drawn from a population with standard deviation $6$. Give the estimate of the population mean, and the standard error of that estimate.
Estimate: m. Standard error: p.
A sample of $70$ items is drawn. Match each statistic computed from that sample to the population quantity it estimates.
| The population mean | The population proportion faulty | The population variance | The largest value in the population | |
|---|---|---|---|---|
| The mean of the sampled values | ||||
| The fraction of sampled items that are faulty | ||||
| The spread of the sampled values about their mean | ||||
| The largest of the sampled values |
A population has variance $12$. Plot the variance of the sample mean for samples of $1$, $2$, $3$ and $4$ observations.
Plot your answer on the grid:
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A sample of $16$ items is about to be drawn, and its mean will be computed. Which quantity is random before the drawing happens?
You can name the parameter, the estimator and the estimate in a problem, and compute a standard error from a population spread and a sample size. Say in your own words why the sample mean is random before the data arrive and the population mean never is.
9. Your turn: a population with variance $50$, and a sample of $200$, step 3