Back to the on-screen lesson ·
Equating sample moments to population moments, one equation per unknown parameter, and checking that the answer is a value the data allow.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will write a population moment as a function of a parameter, set it equal to the matching sample moment, and solve for the estimate, using a second moment when there are two unknowns. You will also check the answer against the parameter space and against the data that produced it.
The first three lessons judged estimators that were handed to you. This one builds them. The only ingredient is the fact that a sample mean estimates a population mean, which the law of large numbers has already supplied, and the only technique is rearranging one equation.
The $k$th population moment is $E[X^{k}]$, a function of the parameters. The $k$th sample moment is $\frac{1}{n}\sum X_i^{k}$, a number the data supply. The method of moments sets as many of the first equal to as many of the second as there are unknown parameters, and solves.
Suppose a model has one unknown parameter $\theta$. Its population mean $E[X]$ is some function of $\theta$, say $g(\theta)$; the sample supplies $\bar x$. The method of moments estimator solves
$$g(\hat\theta) = \bar x.$$
With two unknown parameters, two equations are needed, and the second moment joins the first:
$$E[X] = \bar x, \qquad E[X^{2}] = \frac{1}{n}\sum x_i^{2}.$$
That is the whole method. Its justification is the law of large numbers: the sample moments converge to the population moments, so an estimator built by equating them is consistent whenever $g$ is continuous and invertible.
Its attraction is that it always works and needs no calculus. Its weaknesses are just as plain. It uses only as much of the data as the moments it employs, so a sample's shape is thrown away; it is not in general the most efficient estimator available; it can produce a value outside the parameter space, such as a negative variance; and, as the uniform case shows, it can produce a value the observed data flatly contradict. It is the first estimator to try and rarely the last.
Another way: story
You are told only the average height of a crowd and asked how tall the tallest person could be. If heights are uniform between zero and some ceiling, the average is half the ceiling, so you double it and report. That is the method of moments in one sentence — and it is also why, if somebody in the crowd is taller than your answer, the method has no way to know.
Another way: steps
What the population mean is, for the families this course uses.
| Model | Population mean | Equation solved | Estimate |
|---|---|---|---|
| One trial, success probability $p$ | $p$ | $p = \bar x$ | $\bar x$ |
| Binomial on $m$ trials | $mp$ | $mp = \bar x$ | $\bar x / m$ |
| Poisson, rate $\lambda$ | $\lambda$ | $\lambda = \bar x$ | $\bar x$ |
| Exponential, rate $\lambda$ | $1/\lambda$ | $1/\lambda = \bar x$ | $1/\bar x$ |
| Uniform on $[0, \theta]$ | $\theta/2$ | $\theta/2 = \bar x$ | $2\bar x$ |
Only the algebra in the third column differs. The method itself is the same sentence every time, which is why it is worth having even where a better estimator exists.
Setting the sample mean equal to the parameter. It is equal to the population mean, which is a function of the parameter and is often not the parameter itself.
Using one equation for two unknowns. One equation cannot determine two numbers; the second moment has to be brought in.
Reporting an impossible answer. A negative variance or a ceiling below an observed value is an answer the method is capable of producing and incapable of noticing. Checking is the author's job.
Expecting it to be unbiased. It is consistent, which is a statement about large samples, and that is a different property from unbiasedness.
Waiting times are exponential with rate $\lambda$, and the sample mean is $4$.
One unknown, so one moment.
The population mean is $1/\lambda$, so $1/\hat\lambda = 4$.
Write the mean in terms of the parameter.
$\hat\lambda = 0.25$ events per unit time.
Solve, and check it is positive.
A sample has mean $10$ and average squared value $116$.
Two unknowns, so two moments.
$\hat\mu = 10$, and $E[X^{2}] = \mu^{2} + \sigma^{2}$ gives $116 = 100 + \hat\sigma^{2}$.
The second moment carries both.
$\hat\sigma^{2} = 16$, which divides by $n$ rather than $n - 1$ and is therefore the biased estimator.
Consistent, not unbiased.
The population mean of a binomial on $8$ trials is $8p$.
Write the mean in terms of the parameter.
So $8\hat p = 6$.
Set it equal to the sample mean.
$\hat p = 0.75$, which lies between zero and one and is therefore a possible value.
Observations are uniform on the interval from $0$ to $\theta$, and the sample mean is $3$. What does the method of moments estimate $\theta$ as?
Answer:
A sample from a uniform on the interval from $0$ to $\theta$ has mean $2$. Separately, $20$ independent trials gave $2$ successes. Give the method of moments estimate of $\theta$, and of the success probability.
Uniform: t. Success probability: p.
The sample mean is $5$. Match each model to the equation the method of moments asks you to solve for its parameter.
| $p = 5$ | $\lambda = 5$ | $\theta / 2 = 5$ | $1 / \lambda = 5$ | |
|---|---|---|---|---|
| A single trial with success probability $p$ | ||||
| A Poisson count with rate $\lambda$ | ||||
| A uniform observation on $[0, \theta]$ | ||||
| An exponential waiting time with rate $\lambda$ |
Three separate samples, each with sample mean $9$. Give the method of moments estimate of the parameter named in each row.
| Estimate | |
|---|---|
| Binomial out of $10$ trials: the success probability | |
| Poisson counts: the rate | |
| Uniform on the interval from $0$ to a parameter: that parameter |
A sample from a normal population has mean $2$ and average squared value $9$. Give the method of moments estimates of the population mean and of the population variance.
Mean: m. Variance: v.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Observations are uniform on the interval from $0$ to $\theta$. The sample mean is $8$, so the method of moments gives $\hat\theta = 16$ — but one of the observations is $17$. What has happened?
You can derive a method of moments estimator for the families this course uses, handle two parameters with two moments, and spot an estimate the observed data contradict. Say in your own words what the method throws away.
9. Your turn: a binomial count out of $8$ trials, with sample mean $6$, step 3