Back to the on-screen lesson ·

Randomized experiments

Random assignment removes selection bias: estimate effects and standard errors, check balance and attrition, and handle noncompliance with the Wald ratio.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to analyze a randomized experiment, compute its standard error, and turn an intention-to-treat effect into an effect on compliers.

2. What you already have

From the first lesson you know that a naive comparison equals the effect on the treated plus selection bias. From the omitted-variable lesson you know that the bias vanishes when the omitted characteristics are unrelated to treatment. From the t-test lesson you can build a confidence interval.

A randomized experiment is the design that makes the bias term zero on purpose. This lesson shows why that works, how to analyze an experiment, and the ways real experiments fall short of the ideal — including the problem of people who do not do what they were assigned to do, whose solution leads straight into the next lesson.

3. Terms to use precisely

TermWhat it means
Randomized controlled trialA study in which a chance device assigns units to treatment or control.
Average treatment effect$E[Y_1 - Y_0]$: the average effect over the whole study population.
BalanceSimilar distributions of pre-treatment characteristics in treated and control groups.
Intention to treat (ITT)The difference in outcomes between the groups as assigned, whatever units actually did.
NoncomplianceUnits not taking the treatment they were assigned, in either direction.
CompliersUnits that take the treatment if and only if they are assigned to it.
AttritionLosing units from the sample before the outcome is measured.

4. Randomization makes selection bias zero

When a coin, lottery or computer decides who is treated, treatment cannot be related to anything about the units — not their motivation, ability, health or income, measured or not. So the treated and control groups have the same distribution of untreated outcomes on average:

$$E[Y_0 \mid D = 1] = E[Y_0 \mid D = 0].$$

The selection-bias term from the first lesson is zero, and the difference in means estimates the average treatment effect:

$$\widehat{ATE} = \bar Y_{treated} - \bar Y_{control}, \qquad se = \sqrt{\frac{s_1^2}{n_1} + \frac{s_0^2}{n_0}}.$$

A regression of the outcome on a treatment dummy gives exactly the same estimate, as the second lesson showed. Adding pre-treatment controls does not change what it estimates, because the controls are unrelated to treatment, but it can shrink the standard error by explaining part of the outcome.

Another way: action

Imagine assigning a thousand people to two groups by coin flip. Any trait you can think of — age, schooling, grit — will be roughly equally common in both groups, even traits nobody recorded. That is the whole trick.

Another way: steps

  1. Check balance: compare pre-treatment characteristics across groups.
  2. Estimate the effect as treated mean minus control mean.
  3. Compute the standard error and a confidence interval.
  4. Check attrition, spillovers and compliance.
  5. With noncompliance, report the ITT and divide it by the take-up difference for the effect on compliers.

5. Why randomization beats controls

Multiple regression removes bias only from the variables it includes. Randomization removes it from every variable at once, including the ones nobody thought to measure. In the language of the omitted-variable lesson, it makes $\delta_1 = 0$ for every omitted characteristic simultaneously.

That is why experiments have become central in economics, from the Tennessee class-size study to the Oregon health insurance lottery to thousands of field experiments in development economics recognized by the 2019 Nobel Prize. When an experiment is feasible and ethical, it gives the most credible answer to a causal question.

The guarantee holds on average across possible randomizations. Any one randomization can produce groups that differ by chance, especially in small samples. That is why researchers report a balance table comparing pre-treatment characteristics, and why a large chance imbalance is handled by controlling for that characteristic rather than by discarding the experiment.

6. Why the standard error has this form

The two group means are computed from different people, so their errors are independent, and the variance of their difference is the sum of their variances. Each mean's variance is its group's variance divided by its size. Adding and taking the square root gives $\sqrt{s_1^2/n_1 + s_0^2/n_0}$.

Two design lessons follow. First, precision depends on both groups: a huge treated group with a tiny control group is wasteful, and with equal variances the smallest standard error for a fixed total comes from equal group sizes. Second, reducing the outcome's variance helps as much as adding people. Measuring the outcome at baseline and controlling for it, or randomizing within similar pairs or blocks, can cut the standard error substantially at no cost in bias.

7. Noncompliance: the offer and the treatment

In most social experiments, researchers can randomize an offer but not the treatment itself. Some people offered training never attend; some in the control group find training elsewhere. The groups as assigned remain comparable, so the difference in outcomes by assignment — the intention-to-treat effect — is still an unbiased estimate of the effect of the offer.

Comparing people who actually trained with those who did not would throw away the randomization: attending is a choice, and selection bias returns. The fix keeps the random comparison and rescales it. If the offer raised the share trained by $0.5$ and raised earnings by $2$, and the offer affected earnings only through training, then those induced to train by the offer — the compliers — must have gained $2/0.5 = 4$.

That ratio, the Wald estimator, is the effect for compliers, not for everyone. People who would always train, or never train, may gain more or less. It is also the simplest instrumental-variables estimate, which the next lesson develops.

8. Attrition, spillovers and other threats

Randomization guarantees comparable groups at the moment of assignment. The comparison can still break afterward.

And experiments have their own limit: an effect estimated in one place, at one scale, with one population, may not transfer to another. That is a question of external validity, and it is answered by replication, not by any single study's standard error.

9. When experiments are not possible

Many of the most important questions in economics cannot be answered with an experiment. Nobody can randomly assign countries to trade agreements, children to parents, or workers to recessions. Some experiments would be unethical, such as randomly denying people a treatment known to help; others would be too expensive or too slow, such as waiting forty years to see how schooling affects retirement savings.

The rest of this unit is about natural experiments and quasi-experimental designs: situations where something close to random assignment happened without a researcher arranging it. A lottery for school places, a policy that applied in one state and not its neighbor, a cutoff score that decided who got a scholarship. In each case the task is to argue that the variation used is as good as random for the comparison being made — to find the experiment hidden in the data.

The experimental ideal is the benchmark for judging all of them. For any observational study, ask: what experiment would answer this question, and how closely does the variation used here imitate it? The closer the imitation, the more a study deserves to be believed; the further, the more its conclusions rest on assumptions a reader should see stated plainly and tested wherever the data allow. Each of the next five lessons ends by asking exactly that question of its own design, so that the comparison with the experiment stays in view throughout.

10. Working an experiment, step by step

A trial randomly offers training. Take-up is $60$ percent among those offered and $10$ percent among the rest. Offered people earn $150$ dollars a month more.

  1. ITT. $150$ dollars: the effect of the offer.
  2. First stage. $0.6 - 0.1 = 0.5$: the offer raised training by 50 points.
  3. Effect on compliers. $150/0.5 = 300$ dollars a month.
  4. Whose effect? People who train only because of the offer.
  5. Check. The complier effect is larger than the ITT, as it must be when take-up is incomplete: the ITT averages the effect over many people the offer did not change.

11. How to check an experiment answer

Three checks.

  1. Is the comparison by assignment? Comparing by treatment actually received reintroduces selection bias.
  2. Did you divide by the take-up difference, not the take-up rate? Control-group take-up counts.
  3. Is the standard error from both groups? Each group's variance over its own size, added.

Then state whose effect the number is: everyone's (ATE, with full compliance), the offer's (ITT) or the compliers' (Wald).

12. In the world: the Oregon health insurance lottery

In 2008 Oregon had money to expand Medicaid to about $10{,}000$ low-income adults but far more applicants, so it held a lottery. Economists saw a rare chance: health insurance, randomly assigned. Winning the lottery raised the probability of having Medicaid by about $25$ percentage points — many winners did not complete the paperwork or turned out to be ineligible.

The researchers compared all lottery winners with all losers (the ITT) and divided by the $0.25$ difference in coverage to get the effect of insurance on those who gained it through the lottery. If winners, say, were $6$ percentage points more likely to report good health, the effect of coverage on compliers would be $6/0.25 = 24$ points.

The study found that Medicaid raised health-care use, reduced financial strain and depression, and improved self-reported health, but did not detectably change blood pressure or cholesterol within two years. Every one of those conclusions rests on the lottery: comparing people who chose to enroll with those who did not would have mixed the effect of insurance with whatever leads people to seek it. The lottery did not need the researchers to guess what that was; the coin had already balanced it. That is the quiet power of a lottery.

13. Comparing takers with non-takers undoes the randomization

The most common mistake is analyzing an experiment by what people did instead of what they were assigned. Taking the treatment is a choice, so that comparison has selection bias like any observational one.

A second mistake is treating a chance imbalance as proof the randomization failed. Imbalance happens by chance; control for it and report it.

A third is reading the complier effect as the effect for everyone. It applies to those whose behavior the assignment changed.

14. An experimental effect

  1. Check the design.

    $\text{random assignment}$

    Selection bias is zero on average.

  2. Read the treated mean.

    $\bar Y_1 = 72$

    Small classes.

  3. Read the control mean.

    $\bar Y_0 = 68$

    Regular classes.

  4. Subtract the two means.

    $72 - 68 = 4$

    The estimated effect.

  5. Name the estimand.

    $\widehat{ATE} = 4$

    Average effect in the study population.

15. A standard error and a test

  1. Read the treated spread.

    $s_1 = 12, \ n_1 = 100$

    Treated group.

  2. Read the control spread.

    $s_0 = 9, \ n_0 = 100$

    Control group.

  3. Compute each variance term.

    $1.44 + 0.81 = 2.25$

    Squared SD over $n$.

  4. Take the square root.

    $\sqrt{2.25} = 1.5$

    Standard error of the difference.

  5. Divide the difference.

    $4.5 \div 1.5 = 3$

    The t statistic.

  6. Build the interval.

    $4.5 \pm 2.94$

    From about $1.6$ to $7.4$.

16. Noncompliance and the Wald ratio

  1. Read the outcome gap by assignment.

    $\text{ITT} = 2.4$

    Offer versus no offer.

  2. Read take-up if offered.

    $0.8$

    Most accepted.

  3. Read take-up if not offered.

    $0.2$

    Some found training anyway.

  4. Subtract the take-up rates.

    $0.8 - 0.2 = 0.6$

    The first stage.

  5. Divide the ITT by it.

    $2.4 \div 0.6 = 4$

    Effect on compliers.

  6. Name the group.

    $\text{compliers}$

    Trained because they were offered.

  7. Compare with a naive comparison.

    $\text{trained vs untrained: biased}$

    Attendance is chosen.

17. Your turn: ITT 1.5, take-up 0.6 if offered and 0.1 if not.

  1. Subtract the take-up rates.

    $0.6 - 0.1 = 0.5$

    The first stage.

  2. Divide the ITT by it.

    $1.5 \div 0.5 = 3$

    Effect on compliers.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Check the size.

18. Guided practice

In a randomized trial, students assigned to small classes average $80$ points on a reading test and students assigned to regular classes average $72$. What is the estimated average effect of a small class?

Answer: points

19. Guided practice

Complete the worked solution: a trial has $100$ treated and $100$ control participants, with standard deviations $6$ and $3$. Find each group's variance term and the sum under the square root.

  1. Compute the treated term.

    $6^2 \div 100 =$ a

    Variance of the treated mean.

  2. Compute the control term.

    $3^2 \div 100 =$ b

    Variance of the control mean.

  3. Add the two terms.

    $\text{treated} + \text{control} =$ c

    The variance of the difference; its root is the standard error.

20. Guided practice

In a trial, the treated group has $n_1 = 100$ and standard deviation $12$; the control group has $n_0 = 100$ and standard deviation $9$. The difference in means is $4.5$. Fill in its standard error and t statistic.

Standard error: se. t statistic: t.

21. Practice

In a trial with imperfect compliance, the outcome is $2.4$ higher in the group assigned to treatment. Take-up is $0.8$ in the assigned group and $0.2$ in the control group. Fill in the intention-to-treat effect, the take-up difference and the effect on compliers.

value
intention-to-treat effect
take-up difference
effect on compliers

22. Practice

Match each problem in a job-training experiment to what it threatens.

attrition: the measured groups are no longer comparablenoncompliance: ITT understates the effect of trainingspillovers: the control outcome is affected by treatmentchance imbalance: control for it, report it
trainees who found jobs are easier to reach for the follow-up survey
many people offered training never attend
trainees take jobs that control-group members would have filled
the offered group happens to have more college graduates

23. Practice

A city randomly offers free job-training vouchers. Of those offered, $70\%$ enroll; of those not offered, $30\%$ find their own way into training. Monthly earnings (hundreds of dollars) are $1.2$ higher, on average, in the offered group. What is the effect of training on the earnings of compliers, in hundreds of dollars?

Answer: hundred dollars a month

24. Somewhere new

In Tennessee's Project STAR, kindergartners and teachers were randomly assigned to small or regular classes within each school. Suppose students in small classes averaged the $53$th percentile on a combined test and students in regular classes the $45$th, with a standard error of $0.5$ percentile points on the difference. What is the t statistic for the class-size effect?

Answer: t statistic

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

In a trial with imperfect compliance, the outcome is $2.4$ higher in the group assigned to treatment. Take-up is $0.8$ in the assigned group and $0.2$ in the control group. Fill in the intention-to-treat effect, the take-up difference and the effect on compliers.

value
intention-to-treat effect
take-up difference
effect on compliers

27. What you can do now

You can analyze an experiment. Explain to someone why comparing people who took a treatment with those who did not throws away the randomization.

Working for the steps left to you

17. Your turn: ITT 1.5, take-up 0.6 if offered and 0.1 if not., step 3

$3 > 1.5$

Larger than the ITT.