Back to the on-screen lesson ·
The addition rule, conditional probability and independence, two-way tables and why an accurate test still cries wolf; expectation, variance, and how both behave under scaling and addition.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you can combine probabilities with the words and, or and given, test independence numerically instead of by intuition, and read a conditional probability off a table of counts — including the case where a very accurate test still produces mostly false alarms. You can also find the expected value and variance of a random variable, and use the rules that make them worth having: expectations always add, and variances add when the variables are independent.
You can find the probability of a simple event by counting. What is new here is combining events — or, and, and the hardest of the three, given — and a table that makes all three visible at once.
Joint probability $P(A \text{ and } B)$: both happen.
Conditional probability $P(B \mid A)$: $B$ happens, given that $A$ did.
Independent: $P(A \text{ and } B) = P(A)P(B)$ — knowing one changes nothing about the other.
Mutually exclusive: they cannot both happen, so $P(A \text{ and } B) = 0$. This is the opposite of independent, not a version of it.
Two-way table: counts cross-classified by two variables, with the row, column and grand totals in the margins.
Three rules cover most of what a probability question asks.
Or. $P(A \text{ or } B) = P(A) + P(B) - P(A \text{ and } B)$. Adding counts the overlap twice, so it is subtracted once. When the two cannot both happen the overlap is zero and the rule looks like plain addition — but that is a special case, and assuming it is the general one is a common error.
Given. $P(B \mid A) = \dfrac{P(A \text{ and } B)}{P(A)}$. Conditioning throws away everything outside $A$ and rescales what is left, which is why the denominator is $P(A)$ and never the whole space.
And. Rearranging gives $P(A \text{ and } B) = P(A)\,P(B \mid A)$, and when $A$ and $B$ are independent that becomes $P(A)P(B)$.
Most questions that look like they need Bayes' theorem need only a table of counts. Imagine a population of 10,000, fill in the four cells, and read the answer off: the conditional probability is one cell over its column total. This is the same arithmetic as the formula, and it is far harder to get backwards.
Another way: picture
Picture two overlapping circles inside a rectangle. $P(A \text{ or } B)$ is the area of both circles, and the lens where they overlap is inside both — so adding the circles counts it twice. Conditioning on $A$ throws the rectangle away and treats circle $A$ as the whole world.
Another way: steps
Take a disease held by $1$ person in $1000$, and a test that is right $99\%$ of the time in both directions. In a population of $100{,}000$:
| Test positive | Test negative | Total | |
|---|---|---|---|
| Ill | 99 | 1 | 100 |
| Well | 999 | 98,901 | 99,900 |
| Total | 1,098 | 98,902 | 100,000 |
A positive result means you are in the first column, which holds $1{,}098$ people — and only $99$ of them are ill. So $P(\text{ill} \mid \text{positive}) = 99/1098 \approx 9\%$, even though the test is right $99\%$ of the time.
Nothing is wrong with the test. There are simply a thousand times more well people to be wrong about, and one per cent of a very large number beats ninety-nine per cent of a very small one.
"Mutually exclusive means independent." They are close to opposites. If $A$ and $B$ cannot both happen, then knowing $A$ happened tells you $B$ certainly did not — which is the strongest possible dependence.
"$P(A \mid B)$ is the same as $P(B \mid A)$." They are usually very different. The probability that somebody who tests positive is ill is not the probability that somebody ill tests positive, and the two can differ by a factor of a hundred.
"A very accurate test rarely gives a false alarm." Not when the condition is rare. A test that is $99\%$ accurate, applied to a disease held by one person in ten thousand, produces about a hundred false positives for every true one — because there are so many more healthy people to be wrong about.
$P(A \text{ or } B) = 0.5 + 0.4 - 0.2 = 0.7$.
Overlap removed once.
$P(B \mid A) = 0.2 / 0.5 = 0.4$, which equals $P(B)$.
Divide by the condition.
Since $P(B \mid A) = P(B)$, the two are independent — and $0.5 \times 0.4 = 0.2$ confirms it.
Two tests, one conclusion.
Of $200$ people, $60$ own a bicycle and cycle to work; $40$ own one and do not.
One row of the table.
$P(\text{cycles} \mid \text{owns}) = 60 / 100 = 0.6$: the denominator is the row total.
Condition on what you were told.
$P(\text{owns and cycles}) = 60 / 200 = 0.3$: a different question, with the grand total underneath.
The denominator is the whole difference.
$P(A \text{ and } B) = 0.6 \times 0.5 = 0.3$.
Rearranged conditional.
If $P(B) = 0.5$ as well, then $P(A)P(B) = 0.3$, which matches — so they are independent.
And then $P(A \text{ or } B) = 0.6 + 0.5 - 0.3 = 0.8$.
$P(A) = 0.3$, $P(B) = 0.4$ and $P(A \text{ and } B) = 0.2$. Find $P(A \text{ or } B)$.
answer
$P(A) = 0.3$ and $P(A \text{ and } B) = 0.1$. Find $P(B \mid A)$.
answer
$P(A) = 0.3$, $P(B) = 0.5$ and $P(A \text{ and } B) = 3/20$. Are $A$ and $B$ independent?
A survey of $100$ people gave the two-way table below. Fill in the missing counts.
| Group | Yes | No | Total |
|---|---|---|---|
| Group 1 | 32 | 60 | |
| Group 2 | 40 | ||
| Total | 50 | 50 | 100 |
You can compute a mean and a standard deviation from data. A random variable has the same two summaries, computed from probabilities instead of from observations — and, unlike data, they obey algebra.
Random variable: a numerical outcome of a chance process.
Expected value $E(X)$ or $\mu$: the probability-weighted average.
Variance $\text{Var}(X)$ or $\sigma^2$: the expected squared distance from the mean.
Standard deviation $\sigma$: its square root, in the units of $X$.
Independent: the value of one says nothing about the value of the other.
A random variable $X$ has an expected value and a variance, and both are computed by weighting with probabilities:
$$E(X) = \sum x\,P(x), \qquad \text{Var}(X) = E(X^2) - [E(X)]^2$$
The second formula is the practical one: weight the squares, then subtract the square of the mean. Doing those two squarings in the wrong order produces a negative variance, which is the built-in check.
What makes random variables worth the notation is that these summaries follow rules that data does not:
That last rule is the foundation of every standard error in the rest of this course. Because variances add while the number of observations grows, the standard deviation of an average shrinks like $\sqrt{n}$ rather than like $n$ — which is why quadrupling a sample halves its error, and no faster.
Another way: picture
Picture two independent quantities as arrows at right angles. Their lengths are the standard deviations, and the spread of the sum is the hypotenuse: $\sqrt{3^2 + 4^2} = 5$, not $7$. Independent variation partly cancels, exactly as perpendicular arrows do not simply add.
Another way: steps
"$E(X)$ is a value $X$ can take." Often it is not. The expected number of heads in three tosses is $1.5$, which never happens. An expectation is a balance point, not a prediction.
"Standard deviations add." They do not, ever. Variances add for independent variables, and the square root is taken at the very end. Adding two standard deviations of $3$ and $4$ gives $7$; the right answer is $5$.
"Subtracting means subtracting the variances." For independent variables, $\text{Var}(X - Y) = \text{Var}(X) + \text{Var}(Y)$. Two sources of variation are still two sources of variation when one is taken away from the other — the spread of a difference is wider than either.
$E(X) = 0 \times \frac14 + 1 \times \frac12 + 2 \times \frac14 = 1$.
Weighted average.
$E(X^2) = 0 \times \frac14 + 1 \times \frac12 + 4 \times \frac14 = 1.5$.
Weight the squares.
$\text{Var}(X) = 1.5 - 1^2 = 0.5$, so $\sigma = \sqrt{0.5} \approx 0.71$.
Subtract, then root.
If $\sigma_X = 3$, then $SD(2X + 7) = 2 \times 3 = 6$: the $7$ does nothing.
Shifts do not spread.
If $\sigma_X = 3$ and $\sigma_Y = 4$ and they are independent, $\text{Var}(X + Y) = 9 + 16 = 25$.
Variances add.
So $SD(X + Y) = 5$ — and $SD(X - Y)$ is also $5$, because subtraction adds the variances too.
Root at the end.
$\text{Var}(X) = 36$ and $\text{Var}(Y) = 64$.
Square first.
$\text{Var}(X + Y) = 100$, so $SD(X + Y) = 10$.
And $SD(3X) = 18$, while $SD(X + 3)$ is still $6$ — one stretches, the other only slides.
$X$ takes the value $1$ with probability $\frac{1}{8}$, $3$ with probability $\frac{2}{8}$ and $5$ with probability $\frac{5}{8}$. Find $E(X)$.
answer
The same $X$ — values $0$, $5$, $10$ with probabilities $\frac{3}{10}$, $\frac{2}{10}$, $\frac{5}{10}$ — has $E(X) = 6$. Find $\text{Var}(X)$.
answer
$X$ has standard deviation $7$. Find the standard deviation of $2X + 16$.
answer
A game costs $3$ to play. One time in ten it pays $8$; otherwise it pays nothing. Give the expected gain per play, and then the expected total over $150$ plays.
per play: e, over 150 plays: t
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Of $150$ people, $20$ are in group 1 and answered yes, and $10$ are in group 2 and answered yes. Given that somebody answered yes, what is the probability they are in group 1?
answer
$X$ and $Y$ are independent, with standard deviations $12$ and $16$. Find the standard deviation of $X + Y$.
answer
You can apply the probability rules and compute with random variables. Without looking: what is the denominator of $P(B \mid A)$, and why is the standard deviation of a sum not the sum of the standard deviations?
9. Your turn: $P(A) = 0.6$, $P(B \mid A) = 0.5$, step 3
20. Your turn: $\sigma_X = 6$, $\sigma_Y = 8$, independent, step 3