Back to the on-screen lesson ·
Potential outcomes, the effect on the treated, and why a naive difference in means equals the effect plus selection bias.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to split a naive comparison into the effect on the treated and selection bias, and say which way the bias leans.
From statistics you can compute a mean, compare two groups and read a scatterplot. From intermediate economics you have met a counterfactual: the outcome that would have happened under a different choice, which is never observed for the same unit at the same time.
This course turns that idea into a toolkit. Econometrics is statistics used to answer economic questions, and most of the questions economists care about are causal: what does a policy, a price or a choice actually do?
| Term | What it means |
|---|---|
| Treatment | The cause being studied, recorded as $D = 1$ if a unit receives it and $D = 0$ if not. |
| Potential outcomes | $Y_1$ and $Y_0$: the outcome a unit would have with and without treatment. |
| Counterfactual | The potential outcome that did not happen and is never observed for that unit. |
| Effect on the treated (ATT) | $E[Y_1 - Y_0 \mid D = 1]$: the average effect among those who were treated. |
| Naive difference | The observed treated mean minus the observed untreated mean. |
| Selection bias | $E[Y_0 \mid D=1] - E[Y_0 \mid D=0]$: how the groups differ without treatment. |
| Identification | An argument that the data, under stated assumptions, reveal the causal quantity. |
Every unit — a worker, a firm, a state — has two potential outcomes: $Y_1$ if it gets the treatment and $Y_0$ if it does not. The causal effect for that unit is $Y_1 - Y_0$. The fundamental problem is that we only ever see one of them.
So we compare groups. The naive comparison takes the treated group's mean and subtracts the untreated group's mean. Add and subtract the treated group's counterfactual mean $E[Y_0 \mid D=1]$ and the naive gap splits in two:
$$\underbrace{E[Y_1 \mid D=1] - E[Y_0 \mid D=0]}_{\text{naive}} = \underbrace{E[Y_1 - Y_0 \mid D=1]}_{\text{effect on the treated}} + \underbrace{E[Y_0 \mid D=1] - E[Y_0 \mid D=0]}_{\text{selection bias}}$$
The first piece is what we want. The second is how differently the two groups would have done with no treatment at all. People choose treatments for reasons, and those reasons usually affect the outcome too, so selection bias is almost never zero in data that people generated by choosing.
Another way: action
Take any headline comparison — college graduates earn more, patients who take a vitamin live longer — and ask: would the two groups have differed even without the treatment? If yes, the headline mixes effect and selection.
Another way: steps
The decomposition is not an assumption; it is algebra. Start from the naive gap $E[Y \mid D=1] - E[Y \mid D=0]$. For the treated, the observed outcome is $Y_1$; for the untreated it is $Y_0$. Now add zero in a useful form: subtract $E[Y_0 \mid D=1]$ and add it back.
Grouping the terms gives $\big(E[Y_1 \mid D=1] - E[Y_0 \mid D=1]\big) + \big(E[Y_0 \mid D=1] - E[Y_0 \mid D=0]\big)$. The first bracket is the average effect for the treated, because it compares the same people in two states. The second compares two different groups in the same untreated state, which is exactly what selection means.
Because it is algebra, it always holds. What varies is whether we can learn the counterfactual term. Everything the rest of this course teaches — randomization, regression with controls, instruments, panels, discontinuities — is a strategy for pinning down $E[Y_0 \mid D=1]$ or for making the selection term zero.
Selection bias appears whenever the reason a unit is treated is related to its untreated outcome. Three patterns come up again and again.
The sign of the bias tells you which way the naive comparison is wrong. Before computing anything, it is worth asking which way you expect it to lean.
Not every empirical question is causal, and the tools differ.
A descriptive question asks what is true in the data: what was the median income, how many firms exported. A predictive question asks what to expect: how many customers will come next month, which loans will default. A good prediction can use any variable that helps, including ones that are consequences rather than causes. A causal question asks what would change if we intervened: would a tax cut raise investment, does a smaller class raise scores.
Mixing them up is the most common error in applied work. A variable can predict an outcome brilliantly and have no causal effect on it; umbrella sales predict rain. The first step of any analysis is to say which kind of question it answers.
The effect on the treated is one of several averages a study can target, and which one matters depends on the decision being made.
The average treatment effect (ATE) averages $Y_1 - Y_0$ over everyone, treated or not. It answers the question a planner asks before a program exists: what would happen if the whole population were treated? The effect on the treated averages only over those who actually took the treatment. It answers a different question: was the program worth it for the people who used it?
The two differ whenever effects vary and people sort into treatment by how much they expect to gain. Workers who choose training may be exactly the ones it helps most, so the effect on them overstates what forcing everyone to train would do. Stating the target average before estimating anything prevents a correct number from being used to answer the wrong question. Most of this course estimates one of these two averages, and each design says which.
Take a job-training example. Trainees earn $62$ dollars an hour on average; non-trainees earn $50$. A follow-up study estimates that trainees would have earned $55$ without training.
The naive comparison overstates the program's effect by five dollars, or about seventy percent. Nothing about the arithmetic is hard. The hard part — and the rest of this course — is where a credible number like $55$ comes from.
Run three checks on every answer.
The third check is the one that catches conceptual errors, and it is worth saying out loud before you look at any data.
College graduates in the United States earn far more than people with only a high school diploma: in recent Census data the gap in median weekly earnings is roughly $70$ percent. That is a naive difference.
How much of it is caused by college? People who go to college differ before they enroll: on average they have higher test scores, more educated parents and different ambitions, all of which would raise their earnings anyway. That is positive selection bias, so the naive gap overstates the return to college.
Economists have attacked the counterfactual in the ways this course will teach: comparing twins with different schooling, using changes in compulsory schooling laws as instruments, and comparing students just above and below admission cutoffs. Most of these studies still find large returns, often around $8$ to $12$ percent more earnings per year of schooling, but the number is a careful estimate of the effect, not a restatement of the raw gap.
Put numbers on it. Suppose graduates earn $1{,}500$ dollars a week and non-graduates $900$, a naive gap of $600$. If graduates would have earned $1{,}100$ without college, the effect is $400$ and selection bias is $200$: a third of the raw gap would have been there anyway. That is the difference between a policy that pays off and one that only seems to.
Across US cities, places with more police officers per resident also have more crime. Read naively, police cause crime. The problem is that cities hire police because they have crime: the treatment is targeted at places whose untreated outcome is worse, so selection bias is large and positive for crime.
Economists have looked for changes in police numbers that did not respond to crime: federal hiring grants handed out on a formula, election-year hiring, and terror alerts that moved officers onto particular streets for reasons unrelated to local crime. Studies built on those changes find that more police reduce crime, especially violent crime, which is the opposite sign from the naive comparison. The raw correlation was not a small error; it had the wrong sign, and the decomposition in this lesson is what explains how that can happen.
The most common mistake is reading a naive difference as the effect of the treatment. It equals the effect only when selection bias is zero, which requires that the two groups would have had the same outcomes without treatment.
A second mistake is thinking a large sample fixes selection bias. More data shrink random noise, but the bias term does not depend on sample size; a million self-selected trainees give a very precise estimate of the wrong number.
A third is assuming the bias always inflates the effect. When treatment is targeted at those worse off, the bias is negative and the naive gap can hide a real benefit.
Name the treatment and outcome.
$D = \text{tutored}, \ Y = \text{score}$
Who is compared with whom.
Read the two observed means.
$\bar Y_{1} = 80, \ \bar Y_{0} = 66$
Tutored and not tutored.
Compute the naive difference.
$80 - 66 = 14$
What the raw data show.
Take the counterfactual from a study.
$E[Y_0 \mid D=1] = 70$
Tutored students without tutoring.
Split into effect and bias.
$80 - 70 = 10, \quad 70 - 66 = 4$
Ten points caused; four points selection.
Name the treatment and outcome.
$D = \text{hospital stay}, \ Y = \text{health}$
Health on a 1–10 scale.
Read the observed means.
$\bar Y_1 = 5, \ \bar Y_0 = 7$
Patients report worse health.
Compute the naive difference.
$5 - 7 = -2$
Hospitals seem to hurt.
State the counterfactual.
$E[Y_0 \mid D=1] = 4$
Patients were sick before they went.
Compute the effect on patients.
$5 - 4 = 1$
The stay helped by one point.
Compute and interpret the bias.
$4 - 7 = -3$
Targeting the sick makes the bias negative.
Name the treatment and outcome.
$D = \text{adopted}, \ Y = \text{output}$
Firms, not people.
Read the observed means.
$\bar Y_1 = 54, \ \bar Y_0 = 52$
Adopters produce a little more.
Compute the naive difference.
$54 - 52 = 2$
A small raw gap.
State the counterfactual.
$E[Y_0 \mid D=1] = 49$
Adopters were weaker firms.
Compute the effect.
$54 - 49 = 5$
The software added five units.
Compute the bias.
$49 - 52 = -3$
Struggling firms were the ones that adopted.
Check the identity.
$5 + (-3) = 2$
Effect plus bias equals the naive gap.
Compute the naive difference.
$48 - 40 = 8$
Raw gap.
Compute the effect on the treated.
$48 - 45 = 3$
Against their own counterfactual.
Compute the selection bias.
Students who used a tutoring app scored $84$ on average; students who did not scored $34$. What is the naive difference in mean scores?
Answer: points
Complete the worked solution: firms that adopted new software have mean output $63$; firms that did not have $49$. Had the adopters not adopted, their mean would have been $57$. Find the naive difference, the selection bias and the effect on the adopters.
Find the naive difference.
$63 - 49 =$ a
Adopters' mean minus non-adopters' mean.
Find the selection bias.
$57 - 49 =$ b
How much better adopters would do anyway.
Subtract to get the effect.
$\text{naive} - \text{bias} =$ c
What the software itself added.
Match each question to the kind of question it is.
| causal: the effect of an intervention | prediction: a forecast, whatever the cause | description: a fact about the data | |
|---|---|---|---|
| Would raising the minimum wage reduce teen employment? | |||
| How many customers will visit next month? | |||
| Does a smaller class raise test scores? | |||
| What was the median household income in Ohio last year? |
People who joined a gym have a mean health score of $71$; people who did not have $60$. Had the joiners never joined, their mean would have been $58$. What is the selection bias in the naive comparison?
Answer: points
Treated mean $80$, untreated mean $66$, and the treated group's counterfactual untreated mean $70$. Fill in the naive difference, the effect on the treated and the selection bias.
| value | |
|---|---|
| naive difference | |
| effect on the treated | |
| selection bias |
Workers who chose a training program earn an average of $35$ dollars an hour afterward; workers who did not earn $30$. A careful follow-up study estimates that the trainees would have earned $36$ without the program. What is the program's average effect on the trainees, in dollars an hour?
Answer: dollars an hour
On a 1–10 health scale, people who spent a night in the hospital last year report a mean of $5$; people who did not report $6$. Doctors estimate that, without the hospital stay, the patients would have reported $4$. What is the average effect of the stay on the patients' health?
Answer: points
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Treated mean $80$, untreated mean $66$, and the treated group's counterfactual untreated mean $70$. Fill in the naive difference, the effect on the treated and the selection bias.
| value | |
|---|---|
| naive difference | |
| effect on the treated | |
| selection bias |
You can decompose a naive comparison. Explain to someone why hospital patients can report worse health even if hospitals help.
17. Your turn: treated mean 48, untreated mean 40, treated counterfactual 45., step 3
$45 - 40 = 5$
And check $3 + 5 = 8$.