Back to the on-screen lesson ·

Omitted variable bias

Short slope = long slope + β₂δ₁: compute the bias, recover the long slope, and sign the bias when the omitted variable was never measured.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute omitted-variable bias, move between the short and long regressions, and sign the bias from a verbal story.

2. What you already have

From the last lesson you know a multiple-regression coefficient holds the other regressors fixed, and that it is still biased by anything left out that moves with the regressor. From the first lesson you know selection bias as a difference between two groups' untreated outcomes.

This lesson makes the bias from a left-out variable into a formula with a size and a sign. It is the single most useful result for reading empirical economics, because it lets you ask of any regression: what might be missing, and which way would it push the answer?

3. Terms to use precisely

TermWhat it means
Long regressionThe regression that includes the variable in question: $y$ on $x_1$ and $x_2$.
Short regressionThe regression that leaves it out: $y$ on $x_1$ only.
Omitted variableA variable that affects $y$ and is related to an included regressor but is not in the model.
Auxiliary regressionThe regression of the omitted variable on the included one, with slope $\delta_1$.
Omitted-variable bias$\beta_2\delta_1$: the difference between the short and long slopes.
Upward biasThe short slope is larger than the long slope: $\beta_2\delta_1 > 0$.
ConfounderAnother name for a variable that affects both the regressor and the outcome.

4. Short = long + effect of the omitted variable × its co-movement

Suppose the true model is $y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + u$, but you regress $y$ on $x_1$ alone. Write $x_2$ as its linear prediction from $x_1$ plus a leftover: $x_2 = \delta_0 + \delta_1 x_1 + v$. Substituting gives

$$y = (\beta_0 + \beta_2\delta_0) + (\beta_1 + \beta_2\delta_1)x_1 + (\beta_2 v + u).$$

The short regression estimates the coefficient on $x_1$ in that equation:

$$\tilde\beta_1 = \beta_1 + \beta_2\delta_1.$$

The bias is the product of two things: how much the omitted variable affects $y$ ($\beta_2$), and how much it moves with the included regressor ($\delta_1$). If either is zero there is no bias. If both are positive, or both negative, the short slope is too large; if they have opposite signs, it is too small.

The same identity holds exactly in any sample for the OLS estimates: the short slope equals the long slope plus the long $\hat\beta_2$ times the auxiliary $\hat\delta_1$.

Another way: action

Before reading any regression, name one variable it leaves out. Then ask two questions: does it raise or lower the outcome, and is it higher or lower when the regressor is higher? Multiply the signs.

Another way: steps

  1. Name the omitted variable $x_2$.
  2. Sign or size $\beta_2$: its effect on $y$, holding $x_1$ fixed.
  3. Sign or size $\delta_1$: the slope of $x_2$ on $x_1$.
  4. Multiply to get the bias $\beta_2\delta_1$.
  5. Correct: long = short − bias, or short = long + bias.

5. Why the formula holds

The short regression attributes to $x_1$ everything that moves with $x_1$. When $x_1$ rises by one unit, $x_2$ tends to rise by $\delta_1$ units, and each unit of $x_2$ raises $y$ by $\beta_2$. So a one-unit rise in $x_1$ comes with a total predicted change in $y$ of $\beta_1$ through its own effect plus $\beta_2\delta_1$ through its companion.

The short regression cannot tell those two paths apart. It reports their sum. The long regression can, because it compares units with the same $x_2$, where the second path is shut off. That is the whole content of "holding fixed", put as arithmetic.

6. Signing the bias when the variable was never measured

Often the omitted variable is something nobody measured: ability, motivation, health. You cannot compute $\beta_2\delta_1$, but you can often sign it.

Take the return to schooling. Ability plausibly raises wages ($\beta_2 > 0$) and is higher among people who stay in school longer ($\delta_1 > 0$). The bias is positive, so the simple regression overstates the return. That tells you the direction of the error even without data on ability.

Signing is a powerful argument in two situations. If the bias would push an estimate toward zero and the estimate is still large, the true effect is even larger. If the bias would push it away from zero, a significant estimate may be entirely bias. Economists routinely argue this way in the discussion sections of papers.

7. When omitting a variable does no harm

The formula also says when leaving something out is harmless. If the omitted variable does not affect $y$ once $x_1$ is held fixed ($\beta_2 = 0$), there is no bias. And if it is unrelated to $x_1$ ($\delta_1 = 0$), there is no bias either, however strongly it affects $y$.

The second case is the logic of experiments. When a treatment is assigned at random, every other characteristic — ability, income, motivation — is unrelated to treatment on average, so $\delta_1 = 0$ for all of them at once, measured or not. A randomized experiment is a way of making every omitted-variable bias zero without having to list the omitted variables.

Leaving out an unrelated variable that does affect $y$ still costs something: it stays in the error, which makes the error variance and the standard errors larger. Such a variable is worth including for precision, not for bias.

8. More than one omitted variable

Real regressions leave out many things at once. The formula extends: the short slope equals the long slope plus the sum of $\beta_j\delta_j$ over every omitted variable $j$, where each $\delta_j$ is now a slope from a regression of that variable on the included ones. Some of those terms may be positive and others negative, so biases can partly cancel.

That is both reassuring and dangerous. It is reassuring because a result can survive one plausible bias if another pushes the other way. It is dangerous because an analyst who controls for the variable with a negative bias, and not the one with a positive bias, can make an estimate worse by adding a control. The rule "more controls are always better" is false: adding a variable changes which omitted variables remain, and the net bias of what remains is what matters.

A practical check used in many papers is to add controls one group at a time — demographics, then family background, then prior test scores — and watch how the coefficient of interest moves. If it barely changes as increasingly detailed controls are added, a large bias from yet another omitted variable would have to be quite unlike everything already controlled for. If it keeps moving, the design is fragile, and the reader should be told so.

9. Omitted variables and selection bias are the same problem

The first lesson split a naive comparison into an effect plus selection bias. When the regressor is a treatment dummy, the omitted-variable formula is the same split in regression language. The short regression on the dummy gives the naive difference in means. The long regression, if it included everything that differs between the groups and affects the outcome, would give the effect. The bias $\beta_2\delta_1$ is the selection bias that comes through the omitted characteristic.

Seeing the two as one problem clarifies what every later design is doing. Randomization makes $\delta_1$ zero for all characteristics at once. Instrumental variables use only the part of the treatment that is unrelated to omitted variables. Fixed effects remove omitted variables that do not change over time. Differences-in-differences removes those that change the same way in both groups. Regression discontinuity compares units so close to a cutoff that the omitted variables barely differ. Each is a different way of making the bias term zero when the variable itself cannot be measured. Keep that list in mind: the rest of this course works through each design in turn, and for each one the question is the same, namely which omitted variables it removes and which it leaves in place.

10. Working a correction, step by step

A short regression of log wages on schooling gives $0.10$. Suppose studies with ability measures suggest ability raises log wages by $0.05$ per unit, holding schooling fixed, and each extra year of schooling goes with $0.4$ more units of ability.

  1. Name the pieces. $\tilde\beta_1 = 0.10$, $\beta_2 = 0.05$, $\delta_1 = 0.4$.
  2. Compute the bias. $0.05 \times 0.4 = 0.02$.
  3. Correct. $0.10 - 0.02 = 0.08$.
  4. Interpret. About $8$ percent per year holding ability fixed, instead of $10$.
  5. Check the direction. Both factors positive, so the short slope was too high, as the correction shows.

11. How to check an omitted-variable answer

Three checks.

  1. Does the sign of the bias match the story? Multiply the two signs before computing, and check your number agrees.
  2. Did you add or subtract correctly? Short = long + bias. Going from short to long subtracts the bias; going from long to short adds it.
  3. Is $\delta_1$ the right regression? It is the omitted variable regressed on the included one, not the other way round.

When the bias is large relative to the estimate — as in the class-size data, where it is about half the short slope — say so plainly: the short regression is answering a different question.

12. In the world: class sizes and English learners

In the California district data from earlier lessons, the simple regression of test scores on the student-teacher ratio gives a slope of about $-2.28$ points per student. Add the percentage of students still learning English, and the class-size slope falls to about $-1.10$. The English-learner coefficient is about $-0.65$ points per percentage point.

The formula explains the change exactly. Districts with one more student per teacher have, on average, about $1.8$ percentage points more English learners, so $\delta_1 \approx 1.8$. The bias is $-0.65 \times 1.8 \approx -1.17$, and $-1.10 - 1.17 = -2.27$, matching the short slope. More than half of the simple class-size effect was really the effect of English learners, who are concentrated in districts with larger classes.

That matters for policy. A school board expecting $2.3$ points per student from hiring teachers would be disappointed if the true effect is closer to $1.1$ — and even $1.1$ could still be biased by other omitted differences, such as family income. The formula does not finish the argument, but it shows exactly where each part of the naive number came from, and which parts a better design must still rule out before a board commits money to it.

13. Not every omitted variable biases the slope

The most common mistake is to think any variable left out of a regression biases it. Only variables that both affect $y$ and move with the included regressor do.

A second mistake is getting $\delta_1$ backward — regressing the included variable on the omitted one. The formula needs the omitted variable on the left.

A third is assuming bias only inflates estimates. When the omitted variable's effect and its co-movement have opposite signs, the short slope is too small, and a large bias can even flip its sign.

14. The size of the bias

  1. Name the omitted variable.

    $x_2 = \text{ability}$

    Left out of the wage regression.

  2. Read its effect on wages.

    $\beta_2 = 0.05$

    Holding schooling fixed.

  3. Read its co-movement.

    $\delta_1 = 0.4$

    Ability per year of schooling.

  4. Multiply the two.

    $0.05 \times 0.4 = 0.02$

    The bias.

  5. State the direction.

    $\text{positive: short slope too high}$

    Both factors positive.

15. From the long regression to the short

  1. Read the long slope.

    $\hat\beta_1 = 2$

    With $x_2$ included.

  2. Read the omitted coefficient.

    $\hat\beta_2 = 3$

    Effect of $x_2$.

  3. Read the auxiliary slope.

    $\hat\delta_1 = 0.5$

    $x_2$ on $x_1$.

  4. Compute the bias.

    $3 \times 0.5 = 1.5$

    Positive.

  5. Add it to the long slope.

    $2 + 1.5 = 3.5$

    What the short regression reports.

  6. Check the gap.

    $3.5 - 2 = 1.5$

    Equals the bias.

16. A bias that flips the sign

  1. Read the long slope.

    $\hat\beta_1 = -0.3$

    The true partial effect is negative.

  2. Read the omitted coefficient.

    $\hat\beta_2 = 0.6$

    The omitted variable raises $y$.

  3. Read the auxiliary slope.

    $\hat\delta_1 = 1.5$

    It rises strongly with $x_1$.

  4. Compute the bias.

    $0.6 \times 1.5 = 0.9$

    Large and positive.

  5. Add it to the long slope.

    $-0.3 + 0.9 = 0.6$

    The short slope is positive.

  6. Compare the signs.

    $-0.3 \text{ vs } 0.6$

    The short regression gets the sign wrong.

  7. Draw the lesson.

    $\text{bias can reverse a conclusion}$

    Not just shrink or inflate it.

17. Your turn: short slope 0.8, omitted beta2 = 0.8, delta = −0.5.

  1. Compute the bias.

    $0.8 \times (-0.5) = -0.4$

    Opposite signs: downward.

  2. Subtract it from the short slope.

    $0.8 - (-0.4) = 1.2$

    The long slope.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Check the direction.

18. Guided practice

An omitted variable raises $y$ by $0.7$ per unit, holding $x_1$ fixed. A regression of the omitted variable on $x_1$ has slope $0.6$. What is the omitted-variable bias in the short-regression slope on $x_1$?

Answer: bias

19. Guided practice

Complete the worked solution: a short regression of log wages on schooling gives $0.14$. Ability, omitted, raises log wages by $0.02$ per unit, and a regression of ability on schooling has slope $0.2$. Find the bias, the long slope, and the bias in percentage points of return.

  1. Multiply beta2 by delta.

    $0.02 \times 0.2 =$ a

    The ability bias.

  2. Subtract it from the short slope.

    $0.14 - \text{bias} =$ b

    The return holding ability fixed.

  3. Express the bias in percentage points.

    $100 \times \text{bias} =$ c

    Points of the return that ability accounted for.

20. Guided practice

In the true model of $y$, the slope on $x_1$ is $0.5$ and the slope on $x_2$ is $-2$. A regression of $x_2$ on $x_1$ has slope $0.25$. If $x_2$ is left out, what slope on $x_1$ will the short regression give, on average?

Answer: short-regression slope

21. Practice

The long-regression slope is $0.5$, the omitted variable's coefficient is $-2$, and $\delta_1 = 0.25$. Fill in the bias and the short-regression slope.

value
omitted-variable bias
short-regression slope

22. Practice

Match each omitted-variable story to the direction of the bias in the short slope.

biased upwardbiased downward
wages on schooling, omitting ability (raises wages, higher with schooling)
scores on class size, omitting income (raises scores, lower with bigger classes)
health on hospital visits, omitting prior illness (lowers health, higher with visits)
crop yield on fertilizer, omitting soil quality (raises yield, lower where more fertilizer is used)

23. Practice

A short regression of $y$ on $x_1$ gives a slope of $0.8$. A variable $x_2$ was left out. Other studies estimate that $x_2$ changes $y$ by $0.8$ per unit, holding $x_1$ fixed, and a regression of $x_2$ on $x_1$ has slope $-0.5$. What slope on $x_1$ would the long regression, with $x_2$ included, give?

Answer: long-regression slope

24. Somewhere new

In California districts, test scores fall by $0.6$ points for each percentage point of students still learning English, holding class size fixed; the class-size coefficient in that long regression is $-1.1$. Districts with one more student per teacher have $1.9$ percentage points more English learners. What class-size slope would a regression that omits English learners give?

Answer: points per student

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

The long-regression slope is $0.5$, the omitted variable's coefficient is $-2$, and $\delta_1 = 0.25$. Fill in the bias and the short-regression slope.

value
omitted-variable bias
short-regression slope

27. What you can do now

You can diagnose a missing variable. Explain to someone why leaving ability out of a wage regression overstates the return to schooling.

Working for the steps left to you

17. Your turn: short slope 0.8, omitted beta2 = 0.8, delta = −0.5., step 3

$0.8 < 1.2$

The short slope was too small.