Back to the on-screen lesson ·
One function whose derivatives at zero are the moments, whose products are sums of independent variables, and which determines the distribution.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will read a mean and a variance off a moment generating function by differentiating it at zero, recognise the standard families from the shape of their generating functions, and multiply two of them to identify the distribution of a sum of independent variables. You will also check that a generating function exists before relying on it.
Finding a mean and a variance means two sums or two integrals, done afresh for every distribution. Finding the distribution of a sum of independent variables means a convolution, which is worse. One function makes both routine.
The $k$-th moment of $X$ is $E[X^{k}]$, so the first is the mean and the second is what the variance is built from. The moment generating function is $M(t) = E[e^{tX}]$, defined wherever that expectation is finite. A convolution is the sum-by-sum calculation that finds the distribution of a sum directly, and is the work this lesson exists to avoid.
The moment generating function of $X$ is $$M_X(t) = E\!\left[e^{tX}\right],$$ defined for those $t$ at which the expectation is finite. When it is finite in an interval around $0$, two things follow.
It generates the moments. Expanding the exponential gives $M_X(t) = 1 + tE[X] + \tfrac{t^{2}}{2}E[X^{2}] + \cdots$, so $M^{(k)}(0) = E[X^{k}]$. Differentiate $k$ times, set $t = 0$, and the $k$-th moment falls out. In particular $M(0) = 1$ always — a free check — $M'(0) = \mu$, and $\operatorname{Var}(X) = M''(0) - M'(0)^{2}$.
It turns sums into products. For independent $X$ and $Y$,
$$M_{X+Y}(t) = E[e^{tX}e^{tY}] = M_X(t) \, M_Y(t),$$ where the middle step is exactly where independence is spent. So the distribution of a sum, which otherwise needs a convolution, becomes a multiplication.
It determines the distribution. Two distributions with the same generating function on an interval around zero are the same distribution. That is what makes the previous paragraph useful: multiply, recognise the answer, and you have identified the sum without ever writing its mass function down.
Not every distribution has one. The expectation may be infinite for every $t \ne 0$ — the Cauchy is the standard example — and then none of this is available, which is why the first step is always to check that it exists.
Another way: table
The standard shelf, worth recognising rather than deriving.
| Distribution | $M(t)$ | Where it is finite |
|---|---|---|
| Bernoulli$(p)$ | $1 - p + pe^{t}$ | everywhere |
| Binomial$(n, p)$ | $(1 - p + pe^{t})^{n}$ | everywhere |
| Poisson$(\lambda)$ | $e^{\lambda(e^{t}-1)}$ | everywhere |
| Exponential$(\lambda)$ | $\lambda/(\lambda - t)$ | $t < \lambda$ |
| Normal$(\mu, \sigma^{2})$ | $e^{\mu t + \sigma^{2}t^{2}/2}$ | everywhere |
Another way: steps
A binomial variable is $n$ independent Bernoulli variables added together, so its generating function is the Bernoulli's raised to the $n$-th power — and the whole of the binomial's behaviour follows from that one sentence.
| Read off the power | Gives |
|---|---|
| $M'(0)$ | $np$ |
| $M''(0) - M'(0)^{2}$ | $np(1-p)$ |
| $M_{X}(t)M_{Y}(t)$ for $\mathrm{Bin}(m,p)$ and $\mathrm{Bin}(n,p)$ | $\mathrm{Bin}(m+n, p)$ |
The last line is the one that would be painful otherwise: adding two binomials with the same $p$ by convolution is a page, and by exponents it is $m + n$. Note that it needs the same $p$ — two binomials with different parameters do not combine, and the generating function shows exactly why: the two bases differ, so the powers will not merge.
Multiplying generating functions of dependent variables. The product rule is exactly as strong as the independence it rests on.
Forgetting to evaluate at zero. $M'(t)$ is a function; the moment is $M'(0)$.
Assuming every distribution has one. Check that it is finite near zero first; some distributions have no moment generating function anywhere except at $t = 0$.
Reading $M(t)$ as a probability. It is an expectation of an exponential, so it is generally larger than one, and it is not a distribution of anything.
$M(t) = 1 - p + pe^{t}$, finite everywhere.
Check first.
$M'(t) = pe^{t}$, so $M'(0) = p$; $M''(0) = p$ too.
Both derivatives are the same here.
So the variance is $p - p^{2} = p(1-p)$.
Two moments, one function.
$X \sim \mathrm{Poisson}(2)$ and $Y \sim \mathrm{Poisson}(3)$, independent.
Independence is what licenses the next step.
$M_{X+Y}(t) = e^{2(e^{t}-1)} e^{3(e^{t}-1)} = e^{5(e^{t}-1)}$.
The exponents add.
That is $\mathrm{Poisson}(5)$, so the sum is Poisson with parameter $5$ — no convolution anywhere.
Recognised off the shelf.
$M'(t) = 4e^{t}M(t)$, so $M'(0) = 4$.
The mean.
$M''(0) = 4 + 16 = 20$.
The second moment.
So the variance is $20 - 16 = 4$, equal to the mean: a Poisson.
A count has moment generating function $M(t) = e^{6(e^{t} - 1)}$. Give $M(0)$, $M'(0)$ and $M''(0)$.
| Value | |
|---|---|
| $M(0)$ | |
| $M'(0)$ | |
| $M''(0)$ |
Put the steps of getting the variance out of a moment generating function into order.
Number the steps in order (write the number in the box):
Match each moment generating function to the distribution it belongs to.
| $1 - p + pe^{t}$ | $(1 - p + pe^{t})^{9}$ | $e^{\lambda(e^{t} - 1)}$ | $\dfrac{\lambda}{\lambda - t}$ for $t < \lambda$ | |
|---|---|---|---|---|
| Bernoulli with parameter $p$ | ||||
| Binomial with $9$ trials and parameter $p$ | ||||
| Poisson with parameter $\lambda$ | ||||
| Exponential with rate $\lambda$ |
A count has moment generating function $M(t) = e^{5(e^{t} - 1)}$. Give its mean and its variance.
Mean: p. Variance: q.
$X$ is Poisson with parameter $5$ and $Y$ is Poisson with parameter $3$, independently. What is the variance of $X + Y$?
Answer:
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
$X$ and $Y$ are random variables with generating functions $M_X$ and $M_Y$. When is the generating function of $X + Y$ equal to $M_X(t) M_Y(t)$?
You can extract moments from a generating function, recognise the standard families by theirs, and identify a sum of independent variables by multiplying. Say in your own words which step of the product rule spends the independence.
9. Your turn: $M(t) = e^{4(e^{t}-1)}$; find the mean and the variance, step 3