Back to the on-screen lesson ·

Instrumental variables and two-stage least squares

Recover a causal slope from an endogenous regressor with an instrument: reduced form over first stage, checked for strength, defended on exclusion.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute an IV estimate from means, covariances or regressions, test instrument strength, and argue the conditions an instrument must meet.

2. What you already have

From the last lesson you know the Wald ratio: when a random offer changes who takes a treatment, the effect on compliers is the outcome gap divided by the take-up gap. From the omitted-variable lesson you know why OLS is biased when the regressor moves with something in the error.

Instrumental variables generalize the Wald ratio. They let a researcher estimate a causal slope from observational data whenever there is some source of variation in the regressor that behaves like a random offer. Finding and defending such a source is one of the central skills of empirical economics.

3. Terms to use precisely

TermWhat it means
Endogenous regressorA regressor correlated with the error, so OLS is biased.
InstrumentA variable $z$ that moves the regressor and is unrelated to the error.
Relevance$Cov(z, x) \ne 0$: the instrument actually shifts the regressor.
Exclusion restrictionThe instrument affects the outcome only through the regressor.
First stageThe regression of the endogenous regressor on the instrument.
Reduced formThe regression of the outcome on the instrument.
Two-stage least squaresRegress $x$ on $z$, then $y$ on the fitted $\hat x$; equals the IV ratio with one instrument.

4. Use only the part of x that the instrument moves

Suppose $y = \beta_0 + \beta_1 x + u$ and $x$ is correlated with $u$ — schooling with ability, say. An instrument $z$ must satisfy two conditions:

Take covariances of both sides of the model with $z$. Since $Cov(z, u) = 0$:

$$Cov(z, y) = \beta_1 Cov(z, x) \quad\Rightarrow\quad \beta_{IV} = \frac{Cov(z, y)}{Cov(z, x)}.$$

Dividing numerator and denominator by $Var(z)$ turns this into a ratio of two regression slopes on the instrument: the reduced form (outcome on instrument) over the first stage (regressor on instrument). With a binary instrument it is exactly the Wald ratio of the last lesson.

Two-stage least squares (2SLS) does the same thing in two regressions: first predict $x$ from $z$, then regress $y$ on the prediction $\hat x$. The prediction contains only instrument-driven variation, which is uncorrelated with $u$.

Another way: action

Think of the instrument as a lever that moves schooling for reasons unrelated to ability. Watch how far wages move when the lever moves schooling; that ratio is the effect of schooling itself.

Another way: steps

  1. Name the endogenous regressor and why OLS is biased.
  2. Propose an instrument and argue relevance and exclusion.
  3. First stage: how much does $z$ move $x$? Check its strength.
  4. Reduced form: how much does $z$ move $y$?
  5. Divide: IV = reduced form ÷ first stage.

5. Why the ratio removes the bias

OLS uses all the variation in $x$, including the part that moves with ability, motivation or anything else in the error. The instrument splits $x$ into two pieces: the part the instrument predicts, and the rest. By assumption, the instrument is unrelated to the error, so the predicted part is too. IV uses only that clean piece.

The reduced form measures how the outcome moves with the instrument. Under the exclusion restriction, that movement happens entirely through $x$, so it equals $\beta_1$ times the first stage. Dividing recovers $\beta_1$.

The price is precision. The clean piece of $x$ is usually a small share of its total variation, so IV standard errors are often several times larger than OLS ones. An IV estimate that is unbiased but very imprecise can be less useful than a slightly biased but precise OLS estimate; judging the trade-off requires thinking about how large the OLS bias is likely to be. When the two estimates are close, the instrument has mostly confirmed OLS, and the precise OLS number can be reported with more confidence. When they differ sharply, the difference itself is evidence about the direction and size of the bias, which is often the most interesting finding of all. Report both, side by side, and let the reader see the comparison for themselves. Hiding either one hides information.

6. The exclusion restriction cannot be tested

Relevance can be checked directly: run the first stage and look at the coefficient. Exclusion cannot. It is a claim about a path that does not exist, and with one instrument there is no statistic that detects a direct effect of $z$ on $y$, because any such effect would simply be absorbed into the IV estimate.

So IV is only as good as the argument for exclusion. Distance to college might raise schooling, but families who live near colleges may also live in cities with better-paid jobs — a direct path to wages. Rainfall might raise farm income, but it might also affect children's health or school attendance directly. Good IV papers spend most of their pages on these arguments: showing the instrument is unrelated to pre-existing characteristics, checking outcomes the instrument should not affect, and explaining why each alternative path is implausible.

7. Weak instruments

When the instrument barely moves the regressor, the first stage is close to zero, and dividing by it amplifies everything in the numerator — including any small violation of exclusion and ordinary sampling noise. Weak-instrument IV estimates are biased toward OLS in finite samples and their standard errors are unreliable.

The standard check is the first-stage F statistic for the instruments. With one instrument it is the square of the first-stage t statistic. A common rule of thumb treats F below about 10 as a warning, and recent work suggests higher thresholds for confident inference. A study with a first-stage F of $4$ should not be believed just because its IV estimate is statistically significant.

8. Whose effect does IV estimate?

When effects differ across people, IV estimates the average effect for the compliers — those whose regressor the instrument actually moved. This is the local average treatment effect, or LATE, from the last lesson.

Different instruments therefore estimate different averages. Distance to college moves schooling for people on the margin of attending, often from families with less money; compulsory-schooling laws move it for people who would otherwise have left school at sixteen. Returns estimated with each can differ without either being wrong. A careful report says who the compliers are likely to be and whether the policy in question would affect the same people. A return estimated among people nudged into one more year of high school says little about the return to a doctorate, and a report should not stretch it that far. Stating the complier population alongside the estimate is part of reporting an IV result honestly, in the same way that stating the controls is part of reporting a regression.

9. Finding instruments

Where do instruments come from? Almost always from institutional detail: rules, lotteries, geography, timing and accidents of policy that shift the regressor for reasons unrelated to the outcome.

The discipline is the same in every case. The researcher must explain the mechanism by which the instrument moves the regressor, show that it does so strongly, and rule out, as convincingly as possible, every other path from the instrument to the outcome. An instrument chosen because it happens to be correlated with the regressor in the data, with no story behind it, is almost never credible.

10. Working an IV estimate, step by step

Living near a college (the instrument) raises schooling by $0.4$ years and raises log wages by $0.04$.

  1. First stage. $0.4$ years: relevance, if precisely estimated.
  2. Reduced form. $0.04$ log points.
  3. Ratio. $0.04/0.4 = 0.10$: ten percent per year of schooling.
  4. Compare with OLS. If OLS gives $0.07$, IV is larger — surprising if ability bias pushes OLS up, and a sign that either measurement error attenuated OLS or the compliers have unusually high returns.
  5. Defend exclusion. Does living near a college affect wages other than through schooling? Controls for region and city size help make the case.

11. How to check an IV answer

Three checks.

  1. Is it reduced form over first stage? Upside down gives the reciprocal.
  2. Are both regressions on the same instrument, over the same sample? Mixing samples breaks the ratio.
  3. Is the first stage strong? Compute $F$ before trusting the ratio.

Then write down the exclusion restriction in words and one way it could fail. If you cannot think of any, you have not thought hard enough; if you can, say why it is unlikely to matter much.

12. In the world: the Vietnam draft lottery

Does military service change later earnings? Comparing veterans with non-veterans is hopeless: who serves depends on health, opportunities and preferences that also shape earnings. But in 1970–1972 the United States ran draft lotteries that assigned numbers to birth dates at random, and men with low numbers were eligible to be drafted.

Joshua Angrist used eligibility as an instrument. Eligibility raised the probability of serving by about $16$ percentage points: the first stage. White men who were eligible earned roughly $435$ dollars a year less in the early 1980s: the reduced form. The IV estimate is about $435/0.16 \approx 2{,}700$ dollars a year less earned by men who served because of the draft, roughly fifteen percent of annual earnings at the time.

The exclusion argument is that a birth date's lottery number could affect later earnings only through military service. Critics raised one alternative channel: some eligible men stayed in school longer to avoid the draft, which would raise their earnings and bias the estimate toward zero. Angrist's analysis, and the debate it started, became a model for how instrumental-variables studies are written and challenged. The lottery supplied the randomness; the argument about channels supplied the credibility. Both were needed for the estimate to be taken seriously.

13. A significant first stage does not validate an instrument

The most common mistake is to treat a strong first stage as proof that an instrument is valid. Relevance is only half the requirement; exclusion cannot be tested and must be argued.

A second mistake is dividing by a weak first stage and trusting the result. Small denominators magnify any problem in the numerator.

A third is reading an IV estimate as the effect for everyone. It is the effect for those whose regressor the instrument moved.

14. IV from covariances

  1. Read the instrument's covariance with y.

    $Cov(z, y) = 6$

    The numerator.

  2. Read its covariance with x.

    $Cov(z, x) = 3$

    The denominator.

  3. Divide the two.

    $6 \div 3 = 2$

    The IV slope.

  4. Compute OLS for comparison.

    $Cov(x, y) \div Var(x) = 10 \div 4 = 2.5$

    Uses all variation in $x$.

  5. Interpret the gap.

    $2.5 > 2$

    OLS overstates, if $z$ is valid.

15. The Wald ratio with a lottery

  1. Name the instrument.

    $z = \text{won the lottery}$

    Random.

  2. Read attendance by lottery result.

    $0.7 \text{ vs } 0.3$

    Winners and losers.

  3. Subtract for the first stage.

    $0.7 - 0.3 = 0.4$

    Relevance.

  4. Read scores by lottery result.

    $52 \text{ vs } 50$

    Winners and losers.

  5. Subtract for the reduced form.

    $52 - 50 = 2$

    Effect of winning.

  6. Divide the two.

    $2 \div 0.4 = 5$

    Effect of attending, for compliers.

16. Checking instrument strength

  1. Run the first stage.

    $\hat\pi_1 = 0.12, \ se = 0.04$

    Regressor on instrument.

  2. Compute the t statistic.

    $0.12 \div 0.04 = 3$

    Coefficient over SE.

  3. Square the t statistic.

    $F = 9$

    One instrument.

  4. Compare with the threshold.

    $9 < 10$

    Borderline weak.

  5. Read the consequence.

    $\text{IV biased toward OLS}$

    In finite samples.

  6. Consider the remedy.

    $\text{stronger instrument or robust inference}$

    Weak-IV methods exist.

  7. Report the strength.

    $\text{state } F \text{ with the estimate}$

    Readers need it.

17. Your turn: first stage 0.5, reduced form 1.5.

  1. Write the ratio.

    $\hat\beta_{IV} = \text{RF} \div \text{FS}$

    Reduced form over first stage.

  2. Divide the two.

    $1.5 \div 0.5 = 3$

    The IV estimate.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Name the group.

18. Guided practice

A first-stage regression of the endogenous regressor on a single instrument gives a coefficient of $0.04$ with a standard error of $0.02$. What is the first-stage F statistic?

Answer: F statistic

19. Guided practice

Complete the worked solution: $Cov(z, x) = 0.9$, $Cov(z, y) = 5.4$, $Cov(x, y) = 24$ and $Var(x) = 3$. Find the IV slope, the OLS slope, and how much larger OLS is.

  1. Divide the two covariances.

    $\hat\beta_{IV} = 5.4 \div 0.9 =$ a

    Uses only instrument-driven variation.

  2. Divide for the OLS slope.

    $\hat\beta_{OLS} = 24 \div 3 =$ b

    Uses all the variation in $x$.

  3. Subtract the IV slope.

    $\hat\beta_{OLS} - \hat\beta_{IV} =$ c

    The OLS bias, if the instrument is valid.

20. Guided practice

People who won a lottery for a charter-school seat score $41$ on average; losers score $35$. $0.6$ of winners attend a charter school, against $0.2$ of losers. Fill in the first stage and the reduced form.

First stage: fs. Reduced form: rf.

21. Practice

In a sample, $Cov(z, y) = 1.5$, $Cov(z, x) = 0.5$, $Cov(x, y) = 4$ and $Var(x) = 1$. Fill in the OLS slope and the IV slope using $z$ as the instrument.

value
OLS slope
IV slope

22. Practice

Rainfall is proposed as an instrument for farm income in a study of how income affects children's schooling. Match each worry to the condition it concerns.

relevanceexclusion restrictionexogeneityweak instrument
rainfall barely changes farm income in irrigated areas
heavy rain closes roads, so children miss school directly
rainier regions are also richer for other reasons
the first-stage F statistic is 4

23. Practice

To estimate the return to schooling, an economist uses distance to the nearest college as an instrument. Living near a college raises schooling by $0.4$ years (first stage) and raises log wages by $0.36$ (reduced form). What is the IV estimate of the return to a year of schooling, in log points?

Answer: log points per year

24. Somewhere new

In the Vietnam-era draft lottery, men whose birth dates drew low numbers were draft-eligible. Eligibility raised the probability of military service by $0.15$, and eligible men later earned $345$ dollars a year less on average. What is the IV estimate of the effect of military service on annual earnings, in dollars?

Answer: dollars a year

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

In a sample, $Cov(z, y) = 4$, $Cov(z, x) = 5$, $Cov(x, y) = 9$ and $Var(x) = 6$. Fill in the OLS slope and the IV slope using $z$ as the instrument.

value
OLS slope
IV slope

27. What you can do now

You can use an instrument. Explain to someone why the exclusion restriction cannot be tested with the data.

Working for the steps left to you

17. Your turn: first stage 0.5, reduced form 1.5., step 3

$\text{compliers}$

Whose $x$ the instrument moved.