Back to the on-screen lesson ·
Scoring bias and variance on one scale, why squaring the bias is what makes the comparison possible, and when a biased estimator wins.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will compute a mean squared error from a bias and a variance, use it to choose between estimators that neither dominates, say why the bias is squared and the variance is not, and tell consistency from unbiasedness.
The last two lessons left an estimator with two separate defects: it may aim at the wrong place, and it may scatter. Nothing so far says which is worse, so nothing so far can choose between a tight estimator that aims badly and a wide one that aims well. This lesson puts them on one scale.
The mean squared error of $\hat\theta$ is $E[(\hat\theta - \theta)^{2}]$, the average squared distance from the target. An estimator is consistent when it converges in probability to the parameter as the sample grows, which happens as soon as both its bias and its variance go to zero. Shrinking an estimate means multiplying it by a factor below one, deliberately trading bias for variance.
Write $\hat\theta$ for an estimator of $\theta$. Its mean squared error is
$$\text{MSE}(\hat\theta) = E\big[(\hat\theta - \theta)^{2}\big],$$
and adding and subtracting $E[\hat\theta]$ inside the bracket splits it exactly:
$$\text{MSE}(\hat\theta) = \big(E[\hat\theta] - \theta\big)^{2} + \operatorname{Var}(\hat\theta) = \text{bias}^{2} + \text{variance}.$$
The cross term vanishes because $E[\hat\theta - E[\hat\theta]] = 0$, so this is an identity rather than an approximation.
Two things follow. The first is that bias and variance are now comparable: a bias of $b$ is worth exactly $b^{2}$ units of variance, so a small bias is cheap and a large one is ruinous. The second is that an unbiased estimator is not automatically the best one. An unbiased estimator with variance $10$ loses to a biased one with bias $1$ and variance $2$, which scores $3$.
Consistency is the long-run version of the same statement. If both the bias and the variance of $\hat\theta_n$ go to zero as $n$ grows, then the mean squared error goes to zero and $\hat\theta_n \to \theta$ in probability. An estimator can be biased at every sample size and still be consistent, which is exactly the case of the sample variance divided by $n$.
Another way: picture
Two dartboards. On the first, the darts are scattered wide all round the bullseye; on the second, they are in a tight group a finger's width off it. Mean squared error is the average squared distance from the centre, so it scores both boards with one number — and the tight group off centre often wins.
Another way: steps
One parameter, four candidate estimators, scored on one scale.
| Estimator | Bias | Variance | Mean squared error |
|---|---|---|---|
| A | 0 | 10 | 10 |
| B | 1 | 2 | 3 |
| C | 2 | 1 | 5 |
| D | 4 | 0 | 16 |
Only $A$ is unbiased, and it is the second worst of the four. $B$ beats it by taking on a bias of one in exchange for a fivefold cut in variance; $D$ shows where the trade stops paying, because squaring turned a bias of four into sixteen. The whole of shrinkage estimation lives between $B$ and $C$.
Adding the bias without squaring it. The two terms are in different units until the bias is squared, and an unsquared bias would let a low bias cancel a high variance.
Insisting on an unbiased estimator. Unbiasedness is one property among several. Preferring it whatever the variance is a rule, not a reason.
Reading a small mean squared error as a small error on this sample. It is an average over samples. Any particular estimate may be far out.
Confusing consistency with unbiasedness. They are independent: an estimator can be unbiased and inconsistent, or biased at every $n$ and consistent.
$A$ is unbiased with variance $9$; $B$ has bias $2$ and variance $1$.
Neither dominates on both counts.
$\text{MSE}(A) = 0 + 9 = 9$ and $\text{MSE}(B) = 4 + 1 = 5$.
Square the bias first.
$B$ wins, and would still win with a bias of anything below $\sqrt{8}$.
The trade has a boundary, and it is computable.
Dividing the sum of squared deviations by $n$ gives bias $-\sigma^{2}/n$ and a variance that also falls like $1/n$.
Biased for every finite $n$.
Both terms go to zero, so the mean squared error does, and the estimator is consistent.
Consistency does not require unbiasedness.
The first scores $9 + 4$.
Square the bias, add the variance.
That is $13$, against $0 + 12 = 12$ for the second.
So the unbiased one wins here — narrowly, and it would lose if its variance were $14$.
An estimator has bias $5$ and variance $8$. What is its mean squared error?
Answer:
Three estimators of the same parameter. For each, give the square of its bias and its mean squared error.
| Square of the bias | Mean squared error | |
|---|---|---|
| Bias $0$, variance $9$ | ||
| Bias $2$, variance $3$ | ||
| Bias $4$, variance $3$ |
The sample mean here has variance $4$ and is unbiased. An analyst reports the sample mean plus $3$ instead. Give the variance of what they report, and its mean squared error.
Variance: v. Mean squared error: m.
Four estimators of one parameter, described by their bias and their variance. Match each to its mean squared error.
| $7$ | $25$ | $32$ | $107$ | |
|---|---|---|---|---|
| Bias $0$, variance $7$ | ||||
| Bias $5$, variance $0$ | ||||
| Bias $5$, variance $7$ | ||||
| Bias $10$, variance $7$ |
The sample mean is unbiased for a parameter of $8$ and has variance $8$. An analyst considers reporting a fixed fraction of it instead. Give the mean squared error of each of the three candidates.
| Mean squared error | |
|---|---|
| Report the sample mean itself | |
| Report half the sample mean | |
| Report zero, whatever the data say |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
An estimator of a parameter has variance $4$ and some bias. How is its mean squared error built from the two?
You can score an estimator on one scale, choose between a tight biased estimator and a wide unbiased one, and say what shrinking an estimate towards zero costs and buys. Say in your own words why insisting on an unbiased estimator is a rule rather than a reason.
9. Your turn: an estimator with bias $3$ and variance $4$, against an unbiased one with variance $12$, step 3