Back to the on-screen lesson ·

Probability and random variables

The addition rule, conditional probability and independence, two-way tables and why an accurate test still cries wolf; expectation, variance, and how both behave under scaling and addition.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you can combine probabilities with the words and, or and given, test independence numerically instead of by intuition, and read a conditional probability off a table of counts — including the case where a very accurate test still produces mostly false alarms. You can also find the expected value and variance of a random variable, and use the rules that make them worth having: expectations always add, and variances add when the variables are independent.

2. What you bring to this

You can find the probability of a simple event by counting. What is new here is combining events — or, and, and the hardest of the three, given — and a table that makes all three visible at once.

3. Words you will need

Joint probability $P(A \text{ and } B)$: both happen.

Conditional probability $P(B \mid A)$: $B$ happens, given that $A$ did.

Independent: $P(A \text{ and } B) = P(A)P(B)$ — knowing one changes nothing about the other.

Mutually exclusive: they cannot both happen, so $P(A \text{ and } B) = 0$. This is the opposite of independent, not a version of it.

Two-way table: counts cross-classified by two variables, with the row, column and grand totals in the margins.

4. Or, and, and given

Three rules cover most of what a probability question asks.

Or. $P(A \text{ or } B) = P(A) + P(B) - P(A \text{ and } B)$. Adding counts the overlap twice, so it is subtracted once. When the two cannot both happen the overlap is zero and the rule looks like plain addition — but that is a special case, and assuming it is the general one is a common error.

Given. $P(B \mid A) = \dfrac{P(A \text{ and } B)}{P(A)}$. Conditioning throws away everything outside $A$ and rescales what is left, which is why the denominator is $P(A)$ and never the whole space.

And. Rearranging gives $P(A \text{ and } B) = P(A)\,P(B \mid A)$, and when $A$ and $B$ are independent that becomes $P(A)P(B)$.

Most questions that look like they need Bayes' theorem need only a table of counts. Imagine a population of 10,000, fill in the four cells, and read the answer off: the conditional probability is one cell over its column total. This is the same arithmetic as the formula, and it is far harder to get backwards.

Another way: picture

Picture two overlapping circles inside a rectangle. $P(A \text{ or } B)$ is the area of both circles, and the lens where they overlap is inside both — so adding the circles counts it twice. Conditioning on $A$ throws the rectangle away and treats circle $A$ as the whole world.

Another way: steps

  1. Write down what you are given and what is being asked, with the words and, or, given.
  2. If counts are available, build the two-way table with its margins.
  3. For or, add and subtract the overlap.
  4. For given, divide by the probability of the condition.
  5. Check independence numerically rather than by intuition.

5. Why an accurate test still cries wolf

Take a disease held by $1$ person in $1000$, and a test that is right $99\%$ of the time in both directions. In a population of $100{,}000$:

Test positiveTest negativeTotal
Ill991100
Well99998,90199,900
Total1,09898,902100,000

A positive result means you are in the first column, which holds $1{,}098$ people — and only $99$ of them are ill. So $P(\text{ill} \mid \text{positive}) = 99/1098 \approx 9\%$, even though the test is right $99\%$ of the time.

Nothing is wrong with the test. There are simply a thousand times more well people to be wrong about, and one per cent of a very large number beats ninety-nine per cent of a very small one.

6. Three things that trip people up

"Mutually exclusive means independent." They are close to opposites. If $A$ and $B$ cannot both happen, then knowing $A$ happened tells you $B$ certainly did not — which is the strongest possible dependence.

"$P(A \mid B)$ is the same as $P(B \mid A)$." They are usually very different. The probability that somebody who tests positive is ill is not the probability that somebody ill tests positive, and the two can differ by a factor of a hundred.

"A very accurate test rarely gives a false alarm." Not when the condition is rare. A test that is $99\%$ accurate, applied to a disease held by one person in ten thousand, produces about a hundred false positives for every true one — because there are so many more healthy people to be wrong about.

7. $P(A) = 0.5$, $P(B) = 0.4$, $P(A \text{ and } B) = 0.2$

  1. $P(A \text{ or } B) = 0.5 + 0.4 - 0.2 = 0.7$.

    Overlap removed once.

  2. $P(B \mid A) = 0.2 / 0.5 = 0.4$, which equals $P(B)$.

    Divide by the condition.

  3. Since $P(B \mid A) = P(B)$, the two are independent — and $0.5 \times 0.4 = 0.2$ confirms it.

    Two tests, one conclusion.

8. Reading a conditional off a table

  1. Of $200$ people, $60$ own a bicycle and cycle to work; $40$ own one and do not.

    One row of the table.

  2. $P(\text{cycles} \mid \text{owns}) = 60 / 100 = 0.6$: the denominator is the row total.

    Condition on what you were told.

  3. $P(\text{owns and cycles}) = 60 / 200 = 0.3$: a different question, with the grand total underneath.

    The denominator is the whole difference.

9. Your turn: $P(A) = 0.6$, $P(B \mid A) = 0.5$

  1. $P(A \text{ and } B) = 0.6 \times 0.5 = 0.3$.

    Rearranged conditional.

  2. If $P(B) = 0.5$ as well, then $P(A)P(B) = 0.3$, which matches — so they are independent.

  3. Your turn: work this step out. Its working is at the end of the packet.

    And then $P(A \text{ or } B) = 0.6 + 0.5 - 0.3 = 0.8$.

10. Guided practice

$P(A) = 0.3$, $P(B) = 0.4$ and $P(A \text{ and } B) = 0.2$. Find $P(A \text{ or } B)$.

answer

11. Guided practice

$P(A) = 0.3$ and $P(A \text{ and } B) = 0.1$. Find $P(B \mid A)$.

answer

12. Practice

$P(A) = 0.3$, $P(B) = 0.5$ and $P(A \text{ and } B) = 3/20$. Are $A$ and $B$ independent?

13. Practice

A survey of $100$ people gave the two-way table below. Fill in the missing counts.

GroupYesNoTotal
Group 13260
Group 240
Total5050100

14. What you bring to this

You can compute a mean and a standard deviation from data. A random variable has the same two summaries, computed from probabilities instead of from observations — and, unlike data, they obey algebra.

15. Words you will need

Random variable: a numerical outcome of a chance process.

Expected value $E(X)$ or $\mu$: the probability-weighted average.

Variance $\text{Var}(X)$ or $\sigma^2$: the expected squared distance from the mean.

Standard deviation $\sigma$: its square root, in the units of $X$.

Independent: the value of one says nothing about the value of the other.

16. Two numbers that obey algebra

A random variable $X$ has an expected value and a variance, and both are computed by weighting with probabilities:

$$E(X) = \sum x\,P(x), \qquad \text{Var}(X) = E(X^2) - [E(X)]^2$$

The second formula is the practical one: weight the squares, then subtract the square of the mean. Doing those two squarings in the wrong order produces a negative variance, which is the built-in check.

What makes random variables worth the notation is that these summaries follow rules that data does not:

That last rule is the foundation of every standard error in the rest of this course. Because variances add while the number of observations grows, the standard deviation of an average shrinks like $\sqrt{n}$ rather than like $n$ — which is why quadrupling a sample halves its error, and no faster.

Another way: picture

Picture two independent quantities as arrows at right angles. Their lengths are the standard deviations, and the spread of the sum is the hypotenuse: $\sqrt{3^2 + 4^2} = 5$, not $7$. Independent variation partly cancels, exactly as perpendicular arrows do not simply add.

Another way: steps

  1. List the values and their probabilities; check they total $1$.
  2. $E(X) = \sum x P(x)$.
  3. $E(X^2) = \sum x^2 P(x)$, then $\text{Var}(X) = E(X^2) - [E(X)]^2$.
  4. For $aX + b$: scale the mean by $a$ and shift by $b$; scale the variance by $a^2$.
  5. For sums: add expectations always, add variances only if independent, and take the square root last.

17. Three things that trip people up

"$E(X)$ is a value $X$ can take." Often it is not. The expected number of heads in three tosses is $1.5$, which never happens. An expectation is a balance point, not a prediction.

"Standard deviations add." They do not, ever. Variances add for independent variables, and the square root is taken at the very end. Adding two standard deviations of $3$ and $4$ gives $7$; the right answer is $5$.

"Subtracting means subtracting the variances." For independent variables, $\text{Var}(X - Y) = \text{Var}(X) + \text{Var}(Y)$. Two sources of variation are still two sources of variation when one is taken away from the other — the spread of a difference is wider than either.

18. $X$ is $0, 1, 2$ with probabilities $\frac{1}{4}, \frac{1}{2}, \frac{1}{4}$

  1. $E(X) = 0 \times \frac14 + 1 \times \frac12 + 2 \times \frac14 = 1$.

    Weighted average.

  2. $E(X^2) = 0 \times \frac14 + 1 \times \frac12 + 4 \times \frac14 = 1.5$.

    Weight the squares.

  3. $\text{Var}(X) = 1.5 - 1^2 = 0.5$, so $\sigma = \sqrt{0.5} \approx 0.71$.

    Subtract, then root.

19. Scaling and adding

  1. If $\sigma_X = 3$, then $SD(2X + 7) = 2 \times 3 = 6$: the $7$ does nothing.

    Shifts do not spread.

  2. If $\sigma_X = 3$ and $\sigma_Y = 4$ and they are independent, $\text{Var}(X + Y) = 9 + 16 = 25$.

    Variances add.

  3. So $SD(X + Y) = 5$ — and $SD(X - Y)$ is also $5$, because subtraction adds the variances too.

    Root at the end.

20. Your turn: $\sigma_X = 6$, $\sigma_Y = 8$, independent

  1. $\text{Var}(X) = 36$ and $\text{Var}(Y) = 64$.

    Square first.

  2. $\text{Var}(X + Y) = 100$, so $SD(X + Y) = 10$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    And $SD(3X) = 18$, while $SD(X + 3)$ is still $6$ — one stretches, the other only slides.

21. Guided practice

$X$ takes the value $1$ with probability $\frac{1}{8}$, $3$ with probability $\frac{2}{8}$ and $5$ with probability $\frac{5}{8}$. Find $E(X)$.

answer

22. Guided practice

The same $X$ — values $0$, $5$, $10$ with probabilities $\frac{3}{10}$, $\frac{2}{10}$, $\frac{5}{10}$ — has $E(X) = 6$. Find $\text{Var}(X)$.

answer

23. Practice

$X$ has standard deviation $7$. Find the standard deviation of $2X + 16$.

answer

24. Somewhere new

A game costs $3$ to play. One time in ten it pays $8$; otherwise it pays nothing. Give the expected gain per play, and then the expected total over $150$ plays.

per play: e, over 150 plays: t

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

Of $150$ people, $20$ are in group 1 and answered yes, and $10$ are in group 2 and answered yes. Given that somebody answered yes, what is the probability they are in group 1?

answer

27. Test question

$X$ and $Y$ are independent, with standard deviations $12$ and $16$. Find the standard deviation of $X + Y$.

answer

28. What you can do now

You can apply the probability rules and compute with random variables. Without looking: what is the denominator of $P(B \mid A)$, and why is the standard deviation of a sum not the sum of the standard deviations?

Working for the steps left to you

9. Your turn: $P(A) = 0.6$, $P(B \mid A) = 0.5$, step 3

20. Your turn: $\sigma_X = 6$, $\sigma_Y = 8$, independent, step 3