Back to the on-screen lesson ·

Dummy variables and interactions

Read dummies against a base group, avoid the dummy trap, let slopes differ with interactions, and see a dummy interaction as a difference of differences.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to read category dummies, compute group slopes and predictions with interactions, and recover an interaction from four group means.

2. What you already have

From the second lesson you know a regression on one dummy gives a difference in group means. From the multiple-regression lessons you know how to read a coefficient holding other regressors fixed, and from the functional-form lesson how to turn a log gap into a percent.

Most economic data have categories: region, industry, education level, marital status. And many questions are about whether a relationship differs between groups: is the return to schooling the same for women and men? Is the union wage premium larger in some industries? This lesson gives the regression tools for both.

3. Terms to use precisely

TermWhat it means
Dummy variableA regressor equal to 1 for units in a category and 0 otherwise.
Base groupThe category with no dummy; every dummy coefficient is a gap from it.
Dummy variable trapIncluding a dummy for every category plus an intercept: perfect collinearity.
Intercept shiftA dummy's effect in a model without interactions: a parallel line for the group.
Interaction termThe product of two regressors, such as $female \times educ$.
Slope shiftAn interaction's coefficient on a dummy times a continuous variable: the group's slope differs by it.
Difference of differences$(m_{11} - m_{10}) - (m_{01} - m_{00})$ for four group means.

4. Dummies shift intercepts; interactions shift slopes

A dummy $D$ in $y = \beta_0 + \delta D + \beta_1 x + u$ gives the $D = 1$ group its own intercept, $\beta_0 + \delta$, and the same slope. So $\delta$ is the gap between the groups at any given $x$.

With several categories, include a dummy for each except one, the base group. Each coefficient is that category's gap from the base, and the gap between two non-base categories is the difference of their coefficients. Including all of them alongside an intercept is the dummy variable trap: the dummies add up to one, the same as the intercept's column, so OLS cannot separate them.

An interaction lets the slope differ too:

$$y = \beta_0 + \delta D + \beta_1 x + \gamma (D \times x) + u.$$

The slope is $\beta_1$ for $D = 0$ and $\beta_1 + \gamma$ for $D = 1$. Interacting two dummies gives four groups, and the interaction coefficient is a difference of differences of the four group means.

Another way: action

Draw two lines on one graph, one for each group. A dummy alone moves one line up or down, keeping them parallel. An interaction tilts one line relative to the other.

Another way: steps

  1. Name the base group — the category without a dummy.
  2. Read each dummy coefficient as a gap from the base.
  3. For two non-base groups, subtract their coefficients.
  4. For a group's slope, add the interaction to the main slope.
  5. For a prediction, switch on every term that applies to the group.

5. Why one category must be left out

Suppose every worker is in exactly one of four regions and you include all four region dummies and an intercept. For every observation, the four dummies add up to one — exactly the column of ones the intercept multiplies. Then there are many combinations of coefficients that give the same fitted values: add any number to the intercept and subtract it from all four dummies, and nothing changes. OLS has no unique answer.

Dropping one dummy removes the problem and fixes the meaning of the rest. Which one to drop is a choice about reading, not about fit: every choice gives the same predictions and the same $R^2$. Choose the base that makes the comparisons you care about easiest — often the largest group, or a natural reference like "no diploma".

6. Why the interaction is a difference of differences

Take two dummies, $married$ and $female$, and their product. The four groups have predicted means:

The marriage gap is $\beta_1$ for men and $\beta_1 + \beta_3$ for women. So $\beta_3$ is how much the marriage gap differs between women and men: $(m_{11} - m_{01}) - (m_{10} - m_{00})$. With four groups and four coefficients the regression fits the four means exactly.

The same arithmetic, with "treated" and "after" in place of "female" and "married", is the difference-in-differences design of a later lesson.

7. Reading a dummy in a log regression

In a log-wage regression, a dummy coefficient $\delta$ is approximately a percent gap: $100\delta$ percent. For larger coefficients, use the exact formula $100(e^{\delta} - 1)$. A female coefficient of $-0.2$ means women's predicted wages are about $18$ percent lower than men's with the same regressors, not $20$ percent.

As always, the gap is conditional on the other regressors. A female coefficient in a regression that controls for occupation compares women and men in the same occupation; one that does not compares all women with all men. Both are useful facts, and they answer different questions.

8. Testing whether groups differ

An interaction answers a question that is easy to state and often asked: is the relationship the same in both groups? The null that the slope is the same for women and men is simply $\gamma = 0$, and a t test on the interaction coefficient answers it.

Sometimes the question is broader: does the whole wage equation differ between the groups — intercept, schooling slope, experience slope, everything? Interact the group dummy with every regressor, and test all the interaction coefficients and the dummy jointly with an F test from the last lesson. This is known as a Chow test. Rejecting says the two groups need separate equations; not rejecting says one pooled equation describes both about as well as two would.

Two cautions apply. An insignificant interaction is not proof that the groups are the same, for the reasons of the t-test lesson: the interval may include large differences. And testing many interactions — by region, by age group, by industry — invites the multiple-testing problem, where one of twenty looks significant by chance. Decide which differences you expect before looking.

9. Ordinal categories: dummies or a single variable?

Some categories have an order: no diploma, high school, college, graduate degree. It is tempting to code them $0, 1, 2, 3$ and enter one variable. That forces the wage gap between each step to be the same, which is rarely true: the gap between high school and college is usually much larger than the gap between no diploma and high school.

A set of dummies lets every step have its own gap, at the cost of a few more coefficients. With enough data that cost is small, and the dummies show the shape of the relationship directly. Coding the categories as one number is a restriction, and like any restriction it can be tested with an F test against the dummy version.

10. Interacting two continuous variables

Interactions are not limited to dummies. The effect of fertilizer on crop yields may depend on rainfall; the effect of experience on wages may depend on schooling. The model $y = \beta_0 + \beta_1 x + \beta_2 z + \beta_3 xz + u$ lets the slope on $x$ change continuously with $z$: it is $\beta_1 + \beta_3 z$.

Reading it follows the quadratic lesson. The coefficient $\beta_1$ is the slope on $x$ when $z = 0$, which may be far outside the data — zero rainfall, zero schooling. Report the slope at meaningful values of $z$, such as its mean and one standard deviation either side. Many researchers subtract the mean of $z$ before forming the product, so that $\beta_1$ itself becomes the slope at the average $z$; the fit and the other slopes are unchanged by that choice. Whatever the coding, a table that reports only $\beta_1$ invites readers to misread it, so state the value of $z$ at which each reported slope applies, every single time you report one.

11. Working an interaction, step by step

A log-wage regression gives $0.08$ on $educ$, $-0.2$ on $female$ and $0.01$ on $female \times educ$.

  1. Men's slope. The interaction is off: $0.08$ per year.
  2. Women's slope. $0.08 + 0.01 = 0.09$ per year.
  3. Gap at twelve years. $-0.2 + 0.01 \times 12 = -0.08$: about eight percent.
  4. Gap at sixteen years. $-0.2 + 0.01 \times 16 = -0.04$: about four percent.
  5. Interpret. The gender gap narrows with schooling. The $-0.2$ is the gap at zero years, outside the data, and should not be reported alone.

12. How to check a dummy or interaction answer

Three checks.

  1. Which group is the base? Every dummy coefficient is relative to it.
  2. Is the interaction switched on? Only for units with both characteristics.
  3. Do the four cells reproduce the data? With two interacted dummies, the predicted means must equal the four group means exactly.

And an interpretation check: when a model has an interaction, the main dummy coefficient is the gap at $x = 0$. Report gaps at meaningful values of $x$.

13. In the world: does the payoff to schooling differ for women?

US data show women's wages rise with schooling at least as fast as men's, and often faster. A typical regression of log hourly wages on years of schooling, a female dummy and their interaction might give a slope of $0.09$ for men and an interaction of $0.01$, so women gain about $10$ percent per year of schooling against $9$ for men.

That is one reason women now earn most US bachelor's degrees: the reward to schooling is at least as large for them. It also explains why the average gender gap is smaller among college graduates, conditional on the other regressors. With a female coefficient of $-0.3$, the predicted gap at $12$ years is $-0.3 + 0.01 \times 12 = -0.18$, about $16$ percent, and at $16$ years $-0.3 + 0.16 = -0.14$, about $13$ percent.

Neither number is the effect of discrimination by itself: occupation, hours and career interruptions all differ, and some of those differences themselves reflect constraints women face. The interaction shows how the conditional gap varies with schooling, which is a precise descriptive fact that any explanation has to account for, and a benchmark that later studies with better designs can be measured against when they try to explain it.

14. A main effect is not the whole effect when there is an interaction

The most common mistake is reading the female dummy in a model with $female \times educ$ as the gender gap. It is the gap at zero years of schooling; at any other level, add the interaction times schooling.

A second mistake is forgetting the base group, and reading a region coefficient as the region's wage level rather than its gap from the base.

A third is including a dummy for every category and an intercept. The software will drop one or refuse to run; the coefficients are not identified.

15. Two non-base categories

  1. Name the base group.

    $\text{Northeast}$

    No dummy.

  2. Read the West coefficient.

    $0.05$

    West minus Northeast.

  3. Read the South coefficient.

    $-0.04$

    South minus Northeast.

  4. Subtract the two coefficients.

    $0.05 - (-0.04) = 0.09$

    The base cancels.

  5. Convert to a percent.

    $\approx 9\%$

    West above South, other things fixed.

16. A group's own slope

  1. Write the model.

    $pay = \beta_0 + \delta\,union + 30\,hours + 5\,union \times hours$

    Weekly pay.

  2. Read the non-member slope.

    $30$ per hour

    Interaction off.

  3. Add the interaction.

    $30 + 5 = 35$

    Members' slope.

  4. Compare two members.

    $\Delta hours = 4$

    Same union status.

  5. Multiply by the slope.

    $4 \times 35 = 140$

    Predicted pay gap.

  6. Note the dummy cancels.

    $\delta \text{ drops out}$

    Both are members.

17. The interaction from four means

  1. List the four means.

    $m_{00} = 30, \ m_{10} = 35, \ m_{01} = 27, \ m_{11} = 29$

    Married, female.

  2. Read the intercept.

    $\beta_0 = 30$

    Unmarried men.

  3. Find the marriage gap for men.

    $35 - 30 = 5 = \beta_1$

    Married minus unmarried.

  4. Find the gender gap if unmarried.

    $27 - 30 = -3 = \beta_2$

    Women minus men.

  5. Find the marriage gap for women.

    $29 - 27 = 2$

    Equals $\beta_1 + \beta_3$.

  6. Subtract the two marriage gaps.

    $\beta_3 = 2 - 5 = -3$

    The interaction.

  7. Check the last cell.

    $30 + 5 - 3 - 3 = 29$

    Matches the data.

18. Your turn: slope on educ 0.07, interaction female × educ = −0.01.

  1. Read the men's slope.

    $0.07$

    Interaction off.

  2. Add the interaction.

    $0.07 - 0.01 = 0.06$

    Women's slope.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Convert to a percent.

19. Guided practice

A log-wage regression includes dummies for the South, West and Midwest, with the Northeast as the base group. The coefficients are $South = -0.07$, $West = 0.06$ and $Midwest = -0.03$. By about how many percent do predicted wages in the West exceed those in the South, other regressors held fixed?

Answer: percent

20. Guided practice

Complete the worked solution: a regression of weekly pay on hours, $union$ and $union \times hours$ has a coefficient of $24$ on hours and $8$ on the interaction. Find the slope for union members, the pay gap between two members whose hours differ by $6$, and the same gap for two non-members.

  1. Add the interaction for members.

    $\beta_{hours} + \gamma =$ a

    Members' pay rises faster with hours.

  2. Multiply by the hours gap for members.

    $6 \times \text{member slope} =$ b

    Predicted pay gap between the two members.

  3. Compute the same gap for non-members.

    $6 \times \beta_{hours} =$ c

    The interaction is off for non-members.

21. Guided practice

A log-wage regression includes $educ$, $female$ and $female \times educ$. The coefficients are $0.011$ on $educ$ and $-0.001$ on the interaction. What is the estimated return to a year of schooling for women, as a decimal log-point slope?

Answer: log points per year

22. Practice

A fitted model is $\widehat{wage} = 36 + 7\,married - 5\,female - 5\,married \times female$. What is the predicted wage of a married woman?

Answer: dollars an hour

23. Practice

For $\widehat{wage} = 23 + 5\,married - 2\,female - 5\,married \times female$, fill in the predicted wage for each group.

predicted wage
unmarried men
married men
unmarried women
married women

24. Practice

A wage regression has an intercept and dummies for $HS$, $College$ and $Graduate$ degrees; people with no diploma are the base. Match each quantity to how it is read.

predicted wage with no diplomaCollege wage minus no-diploma wageGraduate wage minus College wageperfect collinearity: the dummy variable trap
the intercept
the College coefficient
Graduate coefficient minus College coefficient
adding a fourth dummy for no diploma

25. Somewhere new

Using national survey data, a researcher estimates $\widehat{\log w} = \beta_0 + 0.011\,educ - 0.2\,female + 0.004\,female \times educ$. By about how many percent do predicted wages rise for a woman who completes $4$ more years of schooling?

Answer: percent

26. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

27. Test question

For $\widehat{wage} = 38 + 8\,married - 5\,female - 3\,married \times female$, fill in the predicted wage for each group.

predicted wage
unmarried men
married men
unmarried women
married women

28. What you can do now

You can read dummies and interactions. Explain to someone why the female coefficient in a model with female × education is not the gender gap.

Working for the steps left to you

18. Your turn: slope on educ 0.07, interaction female × educ = −0.01., step 3

$6\%$ per year

For women.