Back to the on-screen lesson ·
Read dummies against a base group, avoid the dummy trap, let slopes differ with interactions, and see a dummy interaction as a difference of differences.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to read category dummies, compute group slopes and predictions with interactions, and recover an interaction from four group means.
From the second lesson you know a regression on one dummy gives a difference in group means. From the multiple-regression lessons you know how to read a coefficient holding other regressors fixed, and from the functional-form lesson how to turn a log gap into a percent.
Most economic data have categories: region, industry, education level, marital status. And many questions are about whether a relationship differs between groups: is the return to schooling the same for women and men? Is the union wage premium larger in some industries? This lesson gives the regression tools for both.
| Term | What it means |
|---|---|
| Dummy variable | A regressor equal to 1 for units in a category and 0 otherwise. |
| Base group | The category with no dummy; every dummy coefficient is a gap from it. |
| Dummy variable trap | Including a dummy for every category plus an intercept: perfect collinearity. |
| Intercept shift | A dummy's effect in a model without interactions: a parallel line for the group. |
| Interaction term | The product of two regressors, such as $female \times educ$. |
| Slope shift | An interaction's coefficient on a dummy times a continuous variable: the group's slope differs by it. |
| Difference of differences | $(m_{11} - m_{10}) - (m_{01} - m_{00})$ for four group means. |
A dummy $D$ in $y = \beta_0 + \delta D + \beta_1 x + u$ gives the $D = 1$ group its own intercept, $\beta_0 + \delta$, and the same slope. So $\delta$ is the gap between the groups at any given $x$.
With several categories, include a dummy for each except one, the base group. Each coefficient is that category's gap from the base, and the gap between two non-base categories is the difference of their coefficients. Including all of them alongside an intercept is the dummy variable trap: the dummies add up to one, the same as the intercept's column, so OLS cannot separate them.
An interaction lets the slope differ too:
$$y = \beta_0 + \delta D + \beta_1 x + \gamma (D \times x) + u.$$
The slope is $\beta_1$ for $D = 0$ and $\beta_1 + \gamma$ for $D = 1$. Interacting two dummies gives four groups, and the interaction coefficient is a difference of differences of the four group means.
Another way: action
Draw two lines on one graph, one for each group. A dummy alone moves one line up or down, keeping them parallel. An interaction tilts one line relative to the other.
Another way: steps
Suppose every worker is in exactly one of four regions and you include all four region dummies and an intercept. For every observation, the four dummies add up to one — exactly the column of ones the intercept multiplies. Then there are many combinations of coefficients that give the same fitted values: add any number to the intercept and subtract it from all four dummies, and nothing changes. OLS has no unique answer.
Dropping one dummy removes the problem and fixes the meaning of the rest. Which one to drop is a choice about reading, not about fit: every choice gives the same predictions and the same $R^2$. Choose the base that makes the comparisons you care about easiest — often the largest group, or a natural reference like "no diploma".
Take two dummies, $married$ and $female$, and their product. The four groups have predicted means:
The marriage gap is $\beta_1$ for men and $\beta_1 + \beta_3$ for women. So $\beta_3$ is how much the marriage gap differs between women and men: $(m_{11} - m_{01}) - (m_{10} - m_{00})$. With four groups and four coefficients the regression fits the four means exactly.
The same arithmetic, with "treated" and "after" in place of "female" and "married", is the difference-in-differences design of a later lesson.
In a log-wage regression, a dummy coefficient $\delta$ is approximately a percent gap: $100\delta$ percent. For larger coefficients, use the exact formula $100(e^{\delta} - 1)$. A female coefficient of $-0.2$ means women's predicted wages are about $18$ percent lower than men's with the same regressors, not $20$ percent.
As always, the gap is conditional on the other regressors. A female coefficient in a regression that controls for occupation compares women and men in the same occupation; one that does not compares all women with all men. Both are useful facts, and they answer different questions.
An interaction answers a question that is easy to state and often asked: is the relationship the same in both groups? The null that the slope is the same for women and men is simply $\gamma = 0$, and a t test on the interaction coefficient answers it.
Sometimes the question is broader: does the whole wage equation differ between the groups — intercept, schooling slope, experience slope, everything? Interact the group dummy with every regressor, and test all the interaction coefficients and the dummy jointly with an F test from the last lesson. This is known as a Chow test. Rejecting says the two groups need separate equations; not rejecting says one pooled equation describes both about as well as two would.
Two cautions apply. An insignificant interaction is not proof that the groups are the same, for the reasons of the t-test lesson: the interval may include large differences. And testing many interactions — by region, by age group, by industry — invites the multiple-testing problem, where one of twenty looks significant by chance. Decide which differences you expect before looking.
Some categories have an order: no diploma, high school, college, graduate degree. It is tempting to code them $0, 1, 2, 3$ and enter one variable. That forces the wage gap between each step to be the same, which is rarely true: the gap between high school and college is usually much larger than the gap between no diploma and high school.
A set of dummies lets every step have its own gap, at the cost of a few more coefficients. With enough data that cost is small, and the dummies show the shape of the relationship directly. Coding the categories as one number is a restriction, and like any restriction it can be tested with an F test against the dummy version.
Interactions are not limited to dummies. The effect of fertilizer on crop yields may depend on rainfall; the effect of experience on wages may depend on schooling. The model $y = \beta_0 + \beta_1 x + \beta_2 z + \beta_3 xz + u$ lets the slope on $x$ change continuously with $z$: it is $\beta_1 + \beta_3 z$.
Reading it follows the quadratic lesson. The coefficient $\beta_1$ is the slope on $x$ when $z = 0$, which may be far outside the data — zero rainfall, zero schooling. Report the slope at meaningful values of $z$, such as its mean and one standard deviation either side. Many researchers subtract the mean of $z$ before forming the product, so that $\beta_1$ itself becomes the slope at the average $z$; the fit and the other slopes are unchanged by that choice. Whatever the coding, a table that reports only $\beta_1$ invites readers to misread it, so state the value of $z$ at which each reported slope applies, every single time you report one.
A log-wage regression gives $0.08$ on $educ$, $-0.2$ on $female$ and $0.01$ on $female \times educ$.
Three checks.
And an interpretation check: when a model has an interaction, the main dummy coefficient is the gap at $x = 0$. Report gaps at meaningful values of $x$.
US data show women's wages rise with schooling at least as fast as men's, and often faster. A typical regression of log hourly wages on years of schooling, a female dummy and their interaction might give a slope of $0.09$ for men and an interaction of $0.01$, so women gain about $10$ percent per year of schooling against $9$ for men.
That is one reason women now earn most US bachelor's degrees: the reward to schooling is at least as large for them. It also explains why the average gender gap is smaller among college graduates, conditional on the other regressors. With a female coefficient of $-0.3$, the predicted gap at $12$ years is $-0.3 + 0.01 \times 12 = -0.18$, about $16$ percent, and at $16$ years $-0.3 + 0.16 = -0.14$, about $13$ percent.
Neither number is the effect of discrimination by itself: occupation, hours and career interruptions all differ, and some of those differences themselves reflect constraints women face. The interaction shows how the conditional gap varies with schooling, which is a precise descriptive fact that any explanation has to account for, and a benchmark that later studies with better designs can be measured against when they try to explain it.
The most common mistake is reading the female dummy in a model with $female \times educ$ as the gender gap. It is the gap at zero years of schooling; at any other level, add the interaction times schooling.
A second mistake is forgetting the base group, and reading a region coefficient as the region's wage level rather than its gap from the base.
A third is including a dummy for every category and an intercept. The software will drop one or refuse to run; the coefficients are not identified.
Name the base group.
$\text{Northeast}$
No dummy.
Read the West coefficient.
$0.05$
West minus Northeast.
Read the South coefficient.
$-0.04$
South minus Northeast.
Subtract the two coefficients.
$0.05 - (-0.04) = 0.09$
The base cancels.
Convert to a percent.
$\approx 9\%$
West above South, other things fixed.
Write the model.
$pay = \beta_0 + \delta\,union + 30\,hours + 5\,union \times hours$
Weekly pay.
Read the non-member slope.
$30$ per hour
Interaction off.
Add the interaction.
$30 + 5 = 35$
Members' slope.
Compare two members.
$\Delta hours = 4$
Same union status.
Multiply by the slope.
$4 \times 35 = 140$
Predicted pay gap.
Note the dummy cancels.
$\delta \text{ drops out}$
Both are members.
List the four means.
$m_{00} = 30, \ m_{10} = 35, \ m_{01} = 27, \ m_{11} = 29$
Married, female.
Read the intercept.
$\beta_0 = 30$
Unmarried men.
Find the marriage gap for men.
$35 - 30 = 5 = \beta_1$
Married minus unmarried.
Find the gender gap if unmarried.
$27 - 30 = -3 = \beta_2$
Women minus men.
Find the marriage gap for women.
$29 - 27 = 2$
Equals $\beta_1 + \beta_3$.
Subtract the two marriage gaps.
$\beta_3 = 2 - 5 = -3$
The interaction.
Check the last cell.
$30 + 5 - 3 - 3 = 29$
Matches the data.
Read the men's slope.
$0.07$
Interaction off.
Add the interaction.
$0.07 - 0.01 = 0.06$
Women's slope.
Convert to a percent.
A log-wage regression includes dummies for the South, West and Midwest, with the Northeast as the base group. The coefficients are $South = -0.07$, $West = 0.06$ and $Midwest = -0.03$. By about how many percent do predicted wages in the West exceed those in the South, other regressors held fixed?
Answer: percent
Complete the worked solution: a regression of weekly pay on hours, $union$ and $union \times hours$ has a coefficient of $24$ on hours and $8$ on the interaction. Find the slope for union members, the pay gap between two members whose hours differ by $6$, and the same gap for two non-members.
Add the interaction for members.
$\beta_{hours} + \gamma =$ a
Members' pay rises faster with hours.
Multiply by the hours gap for members.
$6 \times \text{member slope} =$ b
Predicted pay gap between the two members.
Compute the same gap for non-members.
$6 \times \beta_{hours} =$ c
The interaction is off for non-members.
A log-wage regression includes $educ$, $female$ and $female \times educ$. The coefficients are $0.011$ on $educ$ and $-0.001$ on the interaction. What is the estimated return to a year of schooling for women, as a decimal log-point slope?
Answer: log points per year
A fitted model is $\widehat{wage} = 36 + 7\,married - 5\,female - 5\,married \times female$. What is the predicted wage of a married woman?
Answer: dollars an hour
For $\widehat{wage} = 23 + 5\,married - 2\,female - 5\,married \times female$, fill in the predicted wage for each group.
| predicted wage | |
|---|---|
| unmarried men | |
| married men | |
| unmarried women | |
| married women |
A wage regression has an intercept and dummies for $HS$, $College$ and $Graduate$ degrees; people with no diploma are the base. Match each quantity to how it is read.
| predicted wage with no diploma | College wage minus no-diploma wage | Graduate wage minus College wage | perfect collinearity: the dummy variable trap | |
|---|---|---|---|---|
| the intercept | ||||
| the College coefficient | ||||
| Graduate coefficient minus College coefficient | ||||
| adding a fourth dummy for no diploma |
Using national survey data, a researcher estimates $\widehat{\log w} = \beta_0 + 0.011\,educ - 0.2\,female + 0.004\,female \times educ$. By about how many percent do predicted wages rise for a woman who completes $4$ more years of schooling?
Answer: percent
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
For $\widehat{wage} = 38 + 8\,married - 5\,female - 3\,married \times female$, fill in the predicted wage for each group.
| predicted wage | |
|---|---|
| unmarried men | |
| married men | |
| unmarried women | |
| married women |
You can read dummies and interactions. Explain to someone why the female coefficient in a model with female × education is not the gender gap.
18. Your turn: slope on educ 0.07, interaction female × educ = −0.01., step 3
$6\%$ per year
For women.