Back to the on-screen lesson ·

Differences-in-differences

Subtract a comparison group's change from the treated group's change, read it as an interaction coefficient, and test the parallel-trends assumption where possible.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute a DiD estimate, build its counterfactual, read the regression form, run a placebo test and name the threats.

2. What you already have

From the dummy-variable lesson you know that the interaction of two dummies equals a difference of differences of four group means. From the panel lesson you know that comparing a unit with itself removes its fixed traits, and that time effects remove shocks common to everyone.

Differences-in-differences combines the two. When a policy reaches one group at a known date, the change in a comparison group over the same period shows what would have happened anyway, and subtracting it from the treated group's change isolates the policy. It is one of the most widely used designs in economics.

3. Terms to use precisely

TermWhat it means
Treated groupUnits exposed to the policy after it starts.
Comparison groupSimilar units not exposed, used to measure what would have happened.
Parallel trendsThe assumption that, without the policy, both groups' outcomes would have changed by the same amount.
CounterfactualThe treated group's outcome had the policy not happened: its start plus the comparison group's change.
Placebo testA DiD computed for periods or groups where no effect should exist.
Event studyA version of DiD that estimates the gap separately for each period before and after the policy.
AnticipationBehavior changing before a policy formally starts, because people expect it.

4. The comparison group's change is the counterfactual

Let $T_0, T_1$ be the treated group's mean before and after, and $C_0, C_1$ the comparison group's. The treated group's change $T_1 - T_0$ mixes the policy's effect with everything else that happened over the period. The comparison group's change $C_1 - C_0$ captures everything else, if the two groups would have moved in parallel. Subtracting leaves the effect:

$$\widehat{DiD} = (T_1 - T_0) - (C_1 - C_0).$$

The same number is the interaction coefficient in

$$y = \beta_0 + \beta_1 Treat + \beta_2 Post + \beta_3\,(Treat \times Post) + u,$$

where $\beta_1$ is the gap between the groups before, $\beta_2$ is the common change, and $\beta_3$ is the DiD. In regression form, controls, many groups and many periods can be added: with group and period fixed effects, it becomes the two-way fixed-effects model of the last lesson.

Differencing within groups removes fixed differences between them, so the groups need not start at the same level. What they need is to have been on the same trend.

Another way: action

Plot both groups' outcomes over time. Before the policy, the two lines should run parallel, even if one is higher. After it, extend the comparison line's slope from the treated group's starting point. The gap between that extension and the treated line is the effect.

Another way: steps

  1. Compute each group's change, after minus before.
  2. Subtract the comparison change from the treated change.
  3. Build the counterfactual: treated before plus comparison change.
  4. Check pre-policy trends with earlier periods or a placebo test.
  5. List spillovers, anticipation and coincident shocks that could break the comparison.

5. Why the double difference works

Write each group's mean as a group level plus a time effect plus the policy effect where it applies: $T_0 = g_T + p_0$, $T_1 = g_T + p_1 + \delta$, $C_0 = g_C + p_0$, $C_1 = g_C + p_1$. The first difference within each group removes its level $g$: $T_1 - T_0 = p_1 - p_0 + \delta$ and $C_1 - C_0 = p_1 - p_0$. The second difference removes the common time change, leaving $\delta$.

The algebra shows exactly what must be true: the time effect must be the same for both groups. If the treated group would have grown faster anyway — because it was richer, younger or recovering from a recession — the difference in time effects is added to the estimate as bias. That is the parallel-trends assumption, and it concerns a counterfactual that can never be observed directly.

6. Checking parallel trends

Parallel trends after the policy cannot be tested, but trends before it can be examined. With several pre-policy periods, plot both groups and look: do the lines run parallel up to the policy date? A placebo test computes a DiD between two pre-policy periods, where the true effect must be zero; a sizable placebo estimate is a warning.

An event study generalizes this. It estimates the treated-minus-comparison gap for each period relative to the policy — three years before, two years before, one year before, the year of, one year after and so on. Flat estimates before the policy and a jump after it are the picture a convincing DiD study shows. Estimates that drift upward before the policy suggest the groups were already diverging.

Passing these checks is reassuring but not proof. A shock that begins at exactly the policy date and affects only the treated group would pass every pre-trend test and still bias the result. That is why a good DiD paper also looks for such shocks directly: it checks news from the treated region around the policy date, examines outcomes the policy should not affect, and shows that the estimated effect appears at the policy date and not before or unrelated to it. Evidence of that kind, gathered from several directions at once, is what persuades a skeptical reader that the comparison group is doing its job.

7. Choosing a comparison group

The credibility of a DiD study rests on its comparison group. Good comparison groups are similar to the treated group in the ways that shape trends: neighboring states, counties just across a border, the same kinds of firms, workers in the same occupations. Card and Krueger compared fast-food restaurants in New Jersey with ones just across the river in eastern Pennsylvania precisely because both faced the same regional economy.

When several comparison groups are plausible, reporting the estimate with each is a useful robustness check. Newer methods build a synthetic control: a weighted average of many untreated units chosen so that its pre-policy path matches the treated unit's as closely as possible. The logic is the same — find a credible stand-in for the counterfactual — with the choice of weights made by the data rather than by hand.

8. Staggered adoption and modern warnings

Many policies roll out at different dates in different places: states adopt a law one by one over fifteen years. The natural extension is a regression with state and year fixed effects and a dummy for whether the law is in force. Recent research has shown this can go wrong when effects grow or shrink over time: the regression then uses already-treated states as comparisons for later-treated ones, and the estimate can even have the wrong sign.

The fix is to compare each newly treated group only with groups not yet treated, and to average those clean comparisons. The two-group, two-period calculation of this lesson is the building block of those methods, and understanding it is what makes the warnings intelligible.

9. Inference with few groups

A DiD comparison of one state with another has, in an important sense, only two units. Every restaurant in New Jersey shares the state's economic shocks, and every restaurant in Pennsylvania shares its state's. Treating four hundred restaurants as four hundred independent observations overstates the precision of the estimate, sometimes enormously: a state-level shock that happens to coincide with the policy looks like a policy effect with a tiny standard error.

The remedy, from the heteroskedasticity lesson, is to cluster standard errors by the unit at which the policy varies — here, the state. With only two states, clustering cannot give reliable standard errors at all, and the honest statement is that the evidence rests on a single comparison. With many treated and comparison states, clustered standard errors work well.

This is one reason modern DiD studies prefer settings with many treated units and many comparison units: dozens of states adopting a policy at different dates, hundreds of counties on either side of many borders. More groups mean more independent comparisons, and precision that reflects the real amount of information in the design rather than the number of rows in the data.

10. Working a DiD estimate, step by step

In February 1992, fast-food restaurants averaged $20.44$ full-time-equivalent workers in New Jersey and $23.33$ in eastern Pennsylvania. By November, after New Jersey raised its minimum wage, the figures were $21.03$ and $21.17$.

  1. New Jersey change. $21.03 - 20.44 = 0.59$.
  2. Pennsylvania change. $21.17 - 23.33 = -2.16$.
  3. DiD. $0.59 - (-2.16) = 2.75$ workers per restaurant.
  4. Counterfactual. $20.44 - 2.16 = 18.28$: New Jersey's employment had it followed Pennsylvania.
  5. Interpret. No evidence that the minimum-wage increase reduced employment; the point estimate is positive.

11. How to check a DiD answer

Three checks.

  1. Are both changes after minus before? Mixing the order flips signs.
  2. Does the counterfactual reproduce the DiD? Treated after minus (treated before plus comparison change) must equal your estimate.
  3. Is the post-period gap between groups the answer? It is not: it includes the groups' original difference in levels.

Then state the parallel-trends assumption in words for this setting, and name one reason it might fail.

12. In the world: the New Jersey minimum wage

On April 1, 1992, New Jersey raised its minimum wage from $4.25$ to $5.05$ dollars an hour; Pennsylvania did not. David Card and Alan Krueger surveyed about $400$ fast-food restaurants in New Jersey and eastern Pennsylvania before and after the increase.

Average full-time-equivalent employment per restaurant went from $20.44$ to $21.03$ in New Jersey and from $23.33$ to $21.17$ in Pennsylvania. The DiD is $0.59 - (-2.16) = 2.75$ workers per restaurant. Rather than the fall in employment that the simplest competitive model predicts, the estimate was positive, though not precisely estimated.

The study transformed the minimum-wage debate and the use of natural experiments in economics, and it was argued over for years. Critics questioned the survey data and whether Pennsylvania was a valid counterfactual; a reanalysis with payroll records found smaller effects in either direction. Card shared the 2021 Nobel Prize partly for this work. The lasting lesson is methodological: a well-chosen comparison group and a clear parallel-trends argument can answer a question that decades of regressions on national data had left unsettled. Later studies using many state increases and the same logic have mostly found small employment effects, which is where the careful reading of this one study pointed.

13. The groups need parallel trends, not equal levels

The most common mistake is thinking the treated and comparison groups must start at the same level. Differencing removes level differences; what matters is that they would have changed in parallel.

A second mistake is reading the post-policy gap between groups as the effect. That gap includes the groups' original difference.

A third is treating a passed pre-trend test as proof. It shows the groups moved together before; a shock at the policy date can still bias the estimate.

14. DiD from four means

  1. Read the treated means.

    $T_0 = 30, \ T_1 = 33.5$

    Before and after.

  2. Read the comparison means.

    $C_0 = 28, \ C_1 = 30$

    Same dates.

  3. Compute the treated change.

    $33.5 - 30 = 3.5$

    Effect plus trend.

  4. Compute the comparison change.

    $30 - 28 = 2$

    Trend alone.

  5. Subtract the changes.

    $3.5 - 2 = 1.5$

    The DiD estimate.

15. The regression coefficients

  1. Match the base cell.

    $\beta_0 = C_0 = 40$

    Control before.

  2. Find the gap before.

    $\beta_1 = T_0 - C_0 = 45 - 40 = 5$

    Level difference.

  3. Find the common change.

    $\beta_2 = C_1 - C_0 = 41 - 40 = 1$

    Time effect.

  4. Find the interaction.

    $\beta_3 = (47.5 - 45) - 1 = 1.5$

    The DiD.

  5. Check the last cell.

    $40 + 5 + 1 + 1.5 = 47.5$

    Equals $T_1$.

  6. Read the effect.

    $\beta_3 = 1.5$

    Policy effect under parallel trends.

16. A decline that is a positive effect

  1. Read the treated means.

    $T_0 = 12.4, \ T_1 = 12.1$

    Treated fell.

  2. Read the comparison means.

    $C_0 = 11.8, \ C_1 = 12.9$

    Comparison rose.

  3. Compute the treated change.

    $12.1 - 12.4 = -0.3$

    A small fall.

  4. Compute the comparison change.

    $12.9 - 11.8 = 1.1$

    What would have happened.

  5. Subtract the changes.

    $-0.3 - 1.1 = -1.4$

    The DiD.

  6. Build the counterfactual.

    $12.4 + 1.1 = 13.5$

    Treated without the policy.

  7. Read the effect.

    $12.1 - 13.5 = -1.4$

    The policy lowered the outcome.

17. Your turn: treated 18.5 to 19.2; comparison 20.1 to 19.6.

  1. Compute the treated change.

    $19.2 - 18.5 = 0.7$

    After minus before.

  2. Compute the comparison change.

    $19.6 - 20.1 = -0.5$

    The trend.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Subtract the changes.

18. Guided practice

After a job-training program starts in one county, average weekly earnings (tens of dollars) go from $44$ to $49$ there and from $44$ to $49$ in a similar county without the program. What is the difference-in-differences estimate?

Answer: tens of dollars

19. Guided practice

Complete the worked solution: the treated group starts at $32$ and ends at $47$; the comparison group rises by $8$ over the same period. Find the treated group's counterfactual after value, its actual change, and the effect.

  1. Add the comparison change to the start.

    $32 + 8 =$ a

    What parallel trends predict without the policy.

  2. Compute the actual change.

    $47 - 32 =$ b

    Treated after minus before.

  3. Subtract the counterfactual from the outcome.

    $47 - \text{counterfactual} =$ c

    The DiD effect.

20. Guided practice

Mean outcomes are: control before $38$, control after $40$, treated before $45$, treated after $48$. In $y = \beta_0 + \beta_1 Treat + \beta_2 Post + \beta_3 Treat \times Post$, fill in $\beta_1$ and $\beta_2$.

β1 (gap before): b1. β2 (control change): b2.

21. Practice

Treated group: $8.2$ before, $9.9$ after. Comparison group: $7.5$ before, $8$ after. Fill in each group's change and the difference-in-differences.

value
treated change
comparison change
difference-in-differences

22. Practice

A city introduces a tax on sugary drinks; a neighboring city does not. Match each worry to the threat it poses to a DiD estimate of the tax's effect on sales.

spillover into the comparison groupnon-parallel pre-trendsanticipationa coincident shock to the treated group
shoppers cross the border to buy drinks in the neighboring city
drink sales in the two cities were already diverging before the tax
stores stocked up heavily in the month before the tax took effect
a new supermarket opened in the taxed city the same month

23. Practice

A state raises its minimum wage; a neighboring state does not. Average full-time-equivalent employment per fast-food restaurant is $12.4$ before and $12.1$ after in the treated state, and $11.8$ before and $12.9$ after in the neighbor. In the regression $y = \beta_0 + \beta_1 Treat + \beta_2 Post + \beta_3 Treat \times Post$, what is $\hat\beta_3$?

Answer: employees per restaurant

24. Somewhere new

Before trusting a DiD study of a state's paid-leave law, a reviewer runs a placebo test on two years before the law. In the treated state employment rose from $48$ to $53$; in the comparison state from $48$ to $52$ (thousands). What is the placebo difference-in-differences?

Answer: thousand jobs

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

Treated group: $20.44$ before, $21.03$ after. Comparison group: $23.33$ before, $21.17$ after. Fill in each group's change and the difference-in-differences.

value
treated change
comparison change
difference-in-differences

27. What you can do now

You can evaluate a policy with DiD. Explain to someone why the two groups need not start at the same level.

Working for the steps left to you

17. Your turn: treated 18.5 to 19.2; comparison 20.1 to 19.6., step 3

$0.7 - (-0.5) = 1.2$

The DiD.