Back to the on-screen lesson ·
The parameter value under which the observed data are least surprising, found by logarithm and derivative where the likelihood is smooth and by its shape where it is not.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will write a likelihood as a product over the observations, take its logarithm, differentiate and solve for the estimate, and check that the stationary point is a maximum. You will also recognise the models whose likelihood has no stationary point, and find their maximum on the boundary instead.
The method of moments used only as much of the data as a moment or two. This lesson uses all of it, through the model's own density or mass function. What is needed from earlier is the product rule for the probability of independent events, and the fact that a differentiable function is stationary at its maximum.
The likelihood $L(\theta)$ is the probability (or density) of the data that were actually observed, read as a function of $\theta$ with the data held fixed. The log-likelihood $\ell(\theta) = \log L(\theta)$ is its logarithm. The maximum likelihood estimate is the value of $\theta$ at which $L$, and therefore $\ell$, is largest.
For independent observations $x_1, \dots, x_n$ from a model with parameter $\theta$,
$$L(\theta) = \prod_{i=1}^{n} f(x_i; \theta), \qquad \ell(\theta) = \sum_{i=1}^{n} \log f(x_i; \theta),$$
and the maximum likelihood estimate $\hat\theta$ is the value maximising either. Note the direction carefully: $L$ is a function of the parameter, with the data fixed. It is not a probability distribution over $\theta$, and it does not integrate to one.
The usual route is calculus: take logs, differentiate, set $\ell'(\theta) = 0$, solve, and confirm with $\ell''(\hat\theta) < 0$ that the stationary point is a maximum. The logarithm changes nothing that matters, being strictly increasing, and changes everything that is convenient, turning a product into a sum.
The route fails exactly when the likelihood is not smooth in the parameter. The uniform on $[0, \theta]$ is the standard case: $L(\theta) = \theta^{-n}$ for $\theta \ge \max x_i$ and $0$ below it, so it is decreasing where it is positive and jumps at the boundary. Its maximum is at $\hat\theta = \max x_i$, and no derivative would ever find it.
Under mild regularity conditions maximum likelihood estimators are consistent, asymptotically normal and asymptotically efficient — no other reasonable estimator beats them for large samples. They are frequently biased in small ones, which is a real cost and a well-understood one.
Another way: story
A bag holds red and blue counters in an unknown proportion. You draw twenty with replacement and get seventeen red. Which proportion makes that result least surprising? Not a half, under which seventeen reds would be a small miracle; the answer is $0.85$, and the likelihood says so by being largest there.
Another way: steps
| Model | Log-likelihood, up to a constant | Estimate |
|---|---|---|
| $n$ trials, $s$ successes | $s\log p + (n - s)\log(1 - p)$ | $s/n$ |
| Poisson counts | $(\sum x_i)\log\lambda - n\lambda$ | $\bar x$ |
| Exponential rate | $n\log\lambda - \lambda\sum x_i$ | $1/\bar x$ |
| Normal, known variance | $-\sum (x_i - \mu)^2 / 2\sigma^{2}$ | $\bar x$ |
| Uniform on $[0, \theta]$ | $-n\log\theta$ for $\theta \ge \max x_i$ | $\max x_i$ |
Four of the five agree with the method of moments; the fifth does not, and it is the one where the method of moments could produce an impossible answer. Where the two methods disagree, the likelihood is using something about the data that the moments threw away.
Reading the likelihood as a probability for the parameter. It is a function of $\theta$ built from fixed data, and it does not integrate to one. Turning it into a probability about $\theta$ requires a prior, which is the last unit of this course.
Differentiating without checking the shape. For the uniform there is no stationary point, and a learner who differentiates anyway finds nothing and concludes there is no estimate.
Assuming the estimate is unbiased. Maximum likelihood is not unbiased in general. The estimate of a normal variance divides by $n$, and the uniform's ceiling estimate is always at or below the truth.
Skipping the second-derivative check. A stationary point can be a minimum, and reporting it would be reporting the parameter under which the data are most surprising.
Counts $x_1, \dots, x_n$ are Poisson with rate $\lambda$: $\ell = (\sum x_i)\log\lambda - n\lambda$ up to a constant.
Logs first.
$\ell'(\lambda) = \frac{\sum x_i}{\lambda} - n = 0$ gives $\hat\lambda = \bar x$.
Differentiate and solve.
$\ell''(\lambda) = -\frac{\sum x_i}{\lambda^{2}} < 0$, so it is a maximum.
The check, done rather than assumed.
Uniform on $[0, \theta]$: $L(\theta) = \theta^{-n}$ when $\theta \ge \max x_i$, and $0$ otherwise.
Zero for any ceiling the data exceed.
Where it is positive it is strictly decreasing in $\theta$, so it is largest at the smallest admissible $\theta$.
No stationary point anywhere.
$\hat\theta = \max x_i$, which is biased low and can never be impossible.
The opposite failing to the method of moments.
The log-likelihood is $12\log p + 18\log(1 - p)$.
Successes and failures.
Setting the derivative to zero gives $\hat p = 12/30$.
The sample proportion.
So $\hat p = 0.4$, and the second derivative is negative there.
$13$ successes in $20$ independent trials. What is the maximum likelihood estimate of the success probability?
Answer:
Three separate samples. The first is $20$ trials with $11$ successes; the second is Poisson counts with sample mean $8$; the third is uniform on the interval from $0$ to an unknown ceiling, with largest observation $26$. Give the maximum likelihood estimate in each row.
| Estimate | |
|---|---|
| Trials: the success probability | |
| Poisson counts: the rate | |
| Uniform: the ceiling |
A sample of $37$ independent observations is to be used to estimate one parameter by maximum likelihood. Put the steps into the order they are carried out.
Number the steps in order (write the number in the box):
$17$ independent exponential waiting times sum to $170$. Give the maximum likelihood estimate of the mean waiting time, and of the rate.
Mean: m. Rate: r.
Four models, each fitted to $30$ observations by maximum likelihood. Match each to the maximum likelihood estimate of its parameter.
| The proportion of trials that succeeded | The sample mean of the counts | One divided by the sample mean of the waiting times | The largest observation in the sample | |
|---|---|---|---|---|
| Independent trials, success probability unknown | ||||
| Poisson counts, rate unknown | ||||
| Exponential waiting times, rate unknown | ||||
| Uniform on $[0, \theta]$, ceiling unknown |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
The likelihood from $45$ independent observations is a product of $45$ factors, and the first move is always to take its logarithm. Why is that allowed?
You can derive a maximum likelihood estimate for the standard families and say where the calculus route does not apply. Say in your own words why the logarithm may be taken without changing the answer.
9. Your turn: $30$ trials give $12$ successes, step 3