Back to the on-screen lesson ·
Separate model fit, causal identification and counterfactual validity in an empirical comparison.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will separate model fit, causal identification and counterfactual validity in an empirical comparison, showing the calculation and stating the assumptions that make the conclusion valid.
Use the supplied definitions and units. Separate an accounting identity, a behavioral assumption, and a normative criterion before drawing conclusions.
| Term | What it means |
|---|---|
| Counterfactual | The outcome under the specified alternative that was not observed for the same target. |
| Estimand | The precisely defined effect or quantity a study aims to identify. |
| Parallel trends | Equality of mean untreated changes across the relevant groups under the design. |
| No anticipation | The pre-treatment outcome is not already affected by the coming intervention. |
| Intention to treat | The effect of assignment to treatment, distinct from treatment receipt under noncompliance. |
| Sensitivity analysis | An explicit comparison showing how conclusions change under altered assumptions. |
An economic model links assumptions to implications. Evidence can test some implications, estimate parameters or help discriminate among mechanisms. But the question must be precise enough to identify the relevant comparison. Predicting next month's sales, measuring a program's effect on participants and choosing a policy for a different population are related tasks that require different forms of evidence.
A causal effect compares an outcome under an intervention with the outcome that would have occurred for the same target under a specified alternative. The alternative is the counterfactual. Observing a participant after treatment reveals only one of those outcomes. Their earlier outcome, a nonparticipant's outcome or a fitted model prediction can serve as a comparison only under additional assumptions.
Suppose a training participant earns fifty units after a course and forty before. The observed increase is ten. It is not automatically the course's effect: earnings might have changed with experience, regional demand or other events even without training. The causal question asks how much of the difference is attributable to the intervention relative to the relevant no-training path.
Define the treatment, population, outcome and time horizon before calculating. An effect on wages among workers who remain employed differs from an effect on total earnings among everyone initially eligible. Conditioning on a post-treatment outcome such as continued employment can change who enters the comparison. A clear estimand specifies what is being averaged rather than allowing available data to silently redefine the question.
Another way: Random assignment changes the comparison, not every problem
Random assignment can make treatment status independent of potential outcomes in the assignment design. Averaging across sufficiently many assigned units then provides a comparison whose uncertainty can be assessed using that design. Randomization balances groups in expectation; it does not guarantee identical observed characteristics in every finite sample. Baseline imbalance can occur by chance without proving that the assignment failed.
An intention-to-treat effect compares outcomes by assigned group. If some assigned participants do not take up the program, this remains the effect of assignment under the design, not automatically the effect of actually receiving treatment. Comparing only people who comply can destroy the original randomized comparison because compliance may relate to motivation, constraints or likely outcomes.
Attrition and spillovers also require attention. If outcomes disappear more often for one treatment group, the observed sample may no longer represent the original assigned groups. If untreated neighbors benefit from a treated person's training, the control condition is affected by the intervention. These are substantive features of the causal comparison, not problems solved by merely increasing the number of decimal places in a mean.
A study can have strong internal validity for its participants and still require assumptions for a different population or scale. A small pilot may use unusually capable staff, have little effect on market wages or recruit people who actively volunteer. Expanding it could change those conditions. Generalizing an estimate requires a mechanism and evidence about the new setting, not just a statistically precise original result.
Another way: Difference-in-differences constructs a specific counterfactual
When treatment is not randomly assigned, a two-group before-and-after design can sometimes use changes in an untreated comparison group to estimate what would have happened to the treated group. Let T 0 and T 1 be treated-group means before and after, and C 0 and C 1 comparison-group means at the same dates. The difference-in-differences estimate is (T 1-T 0)-(C 1-C 0).
The calculation subtracts the comparison group's change from the treated group's change. Equivalently, it predicts the treated group's untreated after outcome as T 0+(C 1-C 0) and subtracts that prediction from T 1. This alternative expression makes the counterfactual visible. Different starting levels are allowed; the crucial restriction concerns the untreated change rather than equality of initial outcomes.
For treated means forty then fifty-five and comparison means thirty then thirty-five, the observed changes are fifteen and five. The constructed untreated outcome for the treated group is forty-five. The estimated effect is ten under the identifying assumptions. The post-period gap of twenty is not the same quantity, because it includes the groups' original ten-unit level difference.
These four means can always produce an arithmetic difference. Interpreting it causally requires that the treated group's mean untreated change would have matched the comparison group's change, along with the design's other requirements. This parallel-trends condition concerns an unobserved post-treatment counterfactual. It does not follow from the subtraction formula and is not directly verified by the four observed numbers.
Another way: Audit timing, spillovers and the untreated trend
No anticipation means the designated pre-treatment outcome is not already affected by knowledge of the impending intervention. If firms hire early because they expect a subsidy next month, the previous month may no longer be an untreated baseline. Treatment timing must reflect when behavior can respond, not only the date a payment is officially recorded.
A comparison group should remain unaffected by the treatment through relevant channels. If a policy moves customers from untreated shops to treated shops, comparison sales can fall because of the same intervention. Subtracting that fall may exaggerate the effect for treated shops relative to a no-policy world. Naming a group control does not establish that it supplies the required untreated trend.
Other events can also create differential trends. Suppose a transport improvement reaches only the treated district at the same time as a training program. The observed additional earnings change can reflect both. An analysis needs an argument or design that distinguishes them. Adding a list of controls is not automatically enough, especially when controls are themselves changed by treatment.
Multiple earlier periods help investigate whether groups previously moved similarly. A visible pre-trend difference is a warning for a simple design. Failure to reject a pre-trend difference, however, does not prove post-period parallel trends: a short noisy series may have little power, and a new differential shock can begin at treatment. Diagnostics inform a substantive argument; they do not turn the identifying assumption into an observed fact.
Another way: Separate uncertainty about a sample from uncertainty about a model
An estimated effect can vary because a study observes a sample rather than every relevant unit. Standard errors and confidence intervals describe this sampling uncertainty under specified assumptions. They do not by themselves include every possible bias from a failed counterfactual, missing outcomes or poor measurement. A narrow interval around a biased comparison is still a narrow interval around a biased comparison.
Suppose a supplied estimate is eight and its standard error is two. Using the explicitly stipulated approximation estimate plus or minus two standard errors gives an interval from four to twelve. This is a rough large-sample interval under an appropriate sampling model, not an exact guarantee that the true effect lies inside this particular realized interval with a newly assigned probability of ninety-five percent.
The assignment and dependence structure matter. If a policy is assigned at the district level, people within a district can share shocks. Treating each person's observation as an independent assignment can overstate precision. An appropriate analysis considers clustering and the number of independent units. More rows in a dataset do not necessarily mean proportionally more independent causal information.
Sensitivity analysis asks how conclusions change when a key assumption is relaxed. If the observed difference-in-differences is ten but an explicitly specified untreated differential trend would have been three, the corresponding adjusted effect is seven. This does not estimate the hidden trend from nothing. It shows exactly which assumed departure changes the result and by how much, making the argument open to scrutiny.
Another way: Fit, identification and policy portability are different tests
A model may predict observed outcomes closely because it has flexible parameters or uses information related to the outcome. Prediction can be useful without identifying an intervention's causal effect. A fitted relationship between sales and advertising does not by itself reveal what sales would do if advertising changed independently of anticipated demand. Firms may advertise more precisely when they expect stronger demand.
Other designs address different sources of variation. A regression-discontinuity design compares units near a treatment threshold under continuity and no-manipulation conditions. An instrumental-variable design needs a relevant source of treatment variation and an exclusion argument, among other assumptions for the chosen estimand. Their names are not substitutes for explaining why the comparison isolates the effect of interest.
With staggered treatment dates and effects that differ across groups or time, a pooled two-way fixed-effects regression need not equal a transparent average of the desired effects. The simple two-group, two-period arithmetic here is deliberately narrower. Extending it requires defining valid comparison groups and weights rather than assuming that adding more dates preserves the interpretation automatically.
The final policy question also includes costs, distribution and behavioral responses at scale. An estimated local earnings increase is not yet a net social benefit calculation. Resources used in the intervention have opportunity costs; gains for participants can interact with outcomes for others. A defensible report states the estimate, its target and uncertainty, its identifying assumptions, and the extra information needed for a broader policy decision.
A district introduces a training program while a comparison district does not. Average monthly earnings in the treated district rise from forty to fifty-five teaching units. In the comparison district they rise from thirty to thirty-five. The analyst uses the same earnings definition and population coverage at both dates, and the four values are stipulated group means rather than noisy estimates for this arithmetic exercise.
The treated change is fifteen and the comparison change five. If the treated district would have experienced the same five-unit increase without the program, its counterfactual after outcome is forty-five. The resulting estimated program effect is ten. The observed after-period difference of twenty includes the original group gap and would answer a different question.
The analyst then investigates a concern: a transport improvement may have raised untreated earnings in the treated district by an additional three units. Under the explicit sensitivity scenario that its untreated trend would exceed the comparison trend by three, the no-program outcome would be forty-eight and the adjusted program effect seven. The calculation does not prove that three is the correct adjustment; it makes the consequence of that assumption visible.
Before publishing a causal conclusion, the analyst would need evidence about the transport timing, anticipation, spillovers, group composition and earlier trends. A larger dataset can improve precision but does not eliminate these design questions. The report therefore separates the observed differences, the counterfactual implied by parallel trends and the alternative sensitivity scenario. It also avoids generalizing the result to a national program whose participants, costs and labor-market responses have not been studied.
A before-after change is not automatically a treatment effect. Parallel trends concerns an unobserved untreated change; similar pre-trends are informative but not proof. Precision does not remove identification bias, and a local effect is not automatically a complete policy welfare result.
Record the same outcome at both dates.
T:40 to 55; C:30 to 35
The population and measurement definitions must be comparable.
Compute the treated group's change.
55-40=15
This is observed growth, not automatically the effect.
Compute the comparison group's change.
35-30=5
Under parallel trends it supplies the untreated change.
Construct the untreated counterfactual.
40+5=45
The treated group's initial level is retained.
Subtract the counterfactual from the treated outcome.
55-45=10
Interpretation is conditional on the identification assumptions.
State the four supplied group means.
T:80 to 74; C:60 to 50
Both observed series decline.
Find the treated change.
74-80=-6
The sign records an observed earnings decline.
Find the comparison change.
50-60=-10
The comparison decline is larger.
Predict the untreated treated-group outcome.
80-10=70
This uses the stated parallel-trends benchmark.
Calculate the relative effect.
74-70=4
A positive effect can mean cushioning a larger counterfactual decline.
Start with the observed comparison.
T:40 to 55; C:30 to 35; DiD 10
The standard counterfactual assumes equal untreated changes.
State the alternative assumption explicitly.
Untreated treated-group trend is 3 above comparison trend
The three-unit departure is supplied, not identified by these four means.
Construct the revised untreated change.
5+3=8
Only the counterfactual changes; observed earnings do not.
Construct the revised counterfactual level.
40+8=48
The initial treated level is still preserved.
Calculate the scenario-adjusted effect.
55-48=7
The departure reduces the standard estimate by three.
Identify the evidence required for a conclusion.
Timing and magnitude of the differential shock need support
Sensitivity arithmetic is not proof that the alternative assumption is true.
T rises 20 to 29 while C rises 12 to 16.
T change 9; C change 4
Both differences use after minus before.
Construct the no-treatment after level under parallel trends.
20+4=24
The comparison trend shifts the treated baseline.
Compute the conditional effect.
Treated mean 40 before and 55 after; comparison 30 before and 35 after. Same units and population definition. Construct both changes, treated untreated-after counterfactual under parallel trends, and DiD. Causality additionally requires no anticipation, no spillovers and an adequate design.
| Your result | |
|---|---|
| Treated change | |
| Comparison change | |
| Treated untreated-after counterfactual | |
| Difference-in-differences |
Treated means 30 then 44; comparison 20 then 24. Under parallel trends, complete the treated change, untreated-after counterfactual and DiD.
Find observed growth in the treated group.
change
Subtract its own before level from its after level.
Add the comparison trend to the treated baseline.
counterfactual
The counterfactual retains the initial difference in levels.
Compare observed and counterfactual after outcomes.
effect
Causal interpretation still requires the stated design assumptions.
Treated means 80 then 74; comparison 60 then 50. Construct both signed changes, the treated counterfactual under parallel trends and DiD. Keep a positive relative effect distinct from an observed increase.
Treated change: v0. Comparison change: v1. Treated untreated-after counterfactual: v2. Difference-in-differences: v3.
Treated means 25 then 30; comparison 20 then 28. Construct both changes, treated untreated-after level under parallel trends and signed DiD.
Treated change: v0. Comparison change: v1. Treated untreated-after counterfactual: v2. Difference-in-differences: v3.
Observed means are T:50 before,66 after and C:30 before,34 after. Two explanations fit these observations equally well. Model A assumes parallel untreated trends. Model B includes an unmeasured transport improvement affecting only T after treatment, adding 5 to its untreated change relative to C. There is no evidence selecting A over B. Construct the treatment effect under each model. Then audit whether these observations alone identify one common causal effect across both models: enter 1 if they do, 0 if they do not. The last field assesses identification, not which policy you prefer.
Effect under model A: v0. Effect under model B: v1. One effect identified: 1 yes, 0 no: v2.
A fictional evaluation records earnings T:50 before,68 after; C:40 before,46 after. Construct both changes, the treated no-program after level under parallel trends and DiD. These supplied means do not themselves verify the trend assumption or justify national extrapolation.
Treated change: v0. Comparison change: v1. Treated untreated-after counterfactual: v2. Difference-in-differences: v3.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A fresh study has T:60 before,78 after and C:40 before,46 after. First construct DiD and its treated untreated-after counterfactual under parallel trends. Then use an explicit sensitivity scenario: the treated untreated trend would have been 4 units above the comparison trend. Produce the scenario-adjusted effect. No assumption is established by these four observed means alone.
DiD: v0. Parallel-trends counterfactual: v1. Scenario-adjusted effect: v2.
Reconstruct a fresh case without the worked solution. Explain which assumption would change its conclusion and which result is only an accounting or model condition.
10. Complete a change comparison, step 3
29-24=5
The counterfactual assumption, not subtraction alone, supplies causal meaning.