Back to the on-screen lesson ·
Conditional means, the law of iterated expectations, and regression as the best linear approximation to the CEF; a dummy's slope is a difference in means.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to compute conditional and overall means, read a dummy-variable regression as a difference in means, and find the regression line when the CEF is linear.
You can compute a mean and a weighted average, and from the last lesson you know that a comparison of group means answers a causal question only when selection bias is zero. In statistics you fitted a least-squares line to a scatterplot.
This lesson says precisely what that line is trying to estimate, before any question of causation. The object is the conditional expectation function, and regression is the best straight-line summary of it.
| Term | What it means |
|---|---|
| Conditional expectation | $E[Y \mid X = x]$: the population mean of $Y$ among units with $X = x$. |
| CEF | The conditional expectation function: the list of conditional means, one for each $x$. |
| CEF error | $e = Y - E[Y \mid X]$; its mean is zero at every value of $X$. |
| Law of iterated expectations | $E[Y] = E\big[E[Y \mid X]\big]$: the overall mean is the weighted average of the conditional means. |
| Population regression | The line $\beta_0 + \beta_1 X$ that minimizes $E[(Y - \beta_0 - \beta_1 X)^2]$. |
| Dummy variable | A regressor that equals 1 for one group and 0 for the other. |
| Regressor | The variable $X$ on the right-hand side; $Y$ is the dependent variable. |
For every value $x$ of a variable $X$, the units with that value have some average outcome, $E[Y \mid X = x]$. Taken together, these averages form the conditional expectation function (CEF). It is the best possible prediction of $Y$ from $X$: no other function of $X$ has a smaller mean squared prediction error.
Any outcome can then be written as its conditional mean plus an error:
$$Y = E[Y \mid X] + e, \qquad E[e \mid X] = 0.$$
That split is not an assumption. It holds by construction, because $e$ is defined as whatever is left over. The error averages zero at every $X$ simply because the mean was subtracted there.
The population regression line $\beta_0 + \beta_1 X$ is the straight line closest to the CEF. When the CEF is itself a straight line, regression recovers it exactly. When it curves, regression gives its best linear approximation. When $X$ is a dummy, the CEF has only two points, a line always fits them exactly, and the slope is simply the difference in the two group means.
Another way: action
Sort a class by hours of sleep, compute the average test score at each number of hours, and plot those averages. The dots are the sample CEF; a regression line through the raw data runs close to them.
Another way: steps
Suppose $30\%$ of workers are in a union and earn $50$ dollars an hour on average, and $70\%$ are not and earn $40$. The overall mean is not the simple average $45$, because the groups are different sizes. Think of $100$ workers: $30$ of them earn a total of $1{,}500$ and $70$ earn $2{,}800$, so the total is $4{,}300$ and the mean is $43$.
That calculation is $0.3 \times 50 + 0.7 \times 40 = 43$, which is the law of iterated expectations: $E[Y] = \sum_x P(X = x)\,E[Y \mid X = x]$. Averaging the conditional means with the right weights must give back the overall mean, because every unit is counted exactly once. The same law is why the CEF error has mean zero overall: it has mean zero within every group, and a weighted average of zeros is zero.
Let $X$ be $1$ for union members and $0$ otherwise. The regression line $\beta_0 + \beta_1 X$ takes only two values: $\beta_0$ at $X = 0$ and $\beta_0 + \beta_1$ at $X = 1$. The line that best fits two conditional means passes through both of them, because two points always lie on a line.
So $\beta_0 = E[Y \mid X = 0]$ and $\beta_0 + \beta_1 = E[Y \mid X = 1]$, which gives $\beta_1 = E[Y \mid X = 1] - E[Y \mid X = 0]$. The regression slope is exactly the naive difference from the last lesson. That is the first sign of a theme that runs through the course: a regression coefficient is only as causal as the comparison it makes. Running a regression does not remove selection bias; with a dummy regressor it computes the same naive gap.
Wages rise quickly with the first years of experience and then flatten; test scores may rise with sleep up to eight hours and then level off. In cases like these the CEF curves, and a straight line cannot match every conditional mean.
Regression still gives the line that is closest to the CEF on average, weighting each value of $X$ by how common it is. Its slope is then a kind of average slope of the CEF. That is often useful, but it can hide important structure: a slope of zero can come from a CEF that rises and then falls. Later lessons show how to let the line bend, with logs, squares and interactions, when the shape matters.
A practical habit follows. Before trusting a single slope, look at the conditional means themselves: bin the regressor into groups, average the outcome in each, and plot the averages. If the dots fall close to a line, one slope is a fair summary. If they bend, rise and fall, or jump at some value, the slope is hiding something the reader of your results deserves to see, and a more flexible specification is worth the extra coefficients.
Suppose you must guess a worker's wage knowing only their years of schooling, and you are penalized by the square of your error. Whatever you guess for workers with twelve years, the guess that makes the average squared error smallest among them is their mean wage: any other number adds the squared distance between it and the mean to every penalty.
Doing that separately at every value of schooling gives exactly the CEF. So no function of $X$, however complicated, predicts $Y$ better in mean squared error. This is why regression, which approximates the CEF with a line, is the natural starting point for prediction as well as for description.
It also shows what the CEF cannot do. It is the best guess given the $X$ you observe. It says nothing about what would happen to $Y$ if you reached in and changed $X$ for a particular person, which is the causal question from the last lesson. The same numbers answer the prediction question well and the causal one only under extra assumptions.
A survey records mean hourly wages of $20$ dollars at $12$ years of schooling, $24$ at $14$ years and $28$ at $16$ years, and no other schooling levels.
Three checks catch most errors.
And a fourth, conceptual check: have you described the answer as a difference in means, not as an effect? The CEF describes the population; causation needs an argument about selection.
The Bureau of Labor Statistics reports that women working full time in the United States earn roughly $83$ cents for every dollar men earn, measured by median weekly earnings. A regression of pay on a dummy for being a woman reproduces exactly that kind of gap: the coefficient is the women's mean minus the men's mean.
That number is a CEF fact about the population, and it is useful: it tells you how different the average paychecks are. It is not by itself the effect of discrimination, because men and women in the data also differ, on average, in hours, occupations and years of experience. Some of those differences may themselves reflect discrimination earlier in life, which is why economists argue about which to hold fixed.
Later lessons add those variables to the regression and show how the coefficient changes. The point of this lesson is the starting place: a dummy coefficient is a difference in conditional means, and interpreting it causally needs an argument, not a regression command.
In the units of the transfer problem: if women's mean weekly pay is $100$ tens of dollars and men's is $120$, the coefficient on the dummy is $-20$, a gap of two hundred dollars a week, or about seventeen percent of men's pay.
The most common mistake is reading a regression slope as what would happen if $X$ were changed. The slope describes how the conditional mean of $Y$ varies with $X$ in the population; it is causal only when the comparison it makes is free of selection.
A second mistake is averaging group means without weights. The overall mean weights each group by its share; the simple average of two group means is right only when the groups are the same size.
A third is treating $E[e \mid X] = 0$ as an assumption about the world. For the CEF error it holds by definition. What is an assumption, later, is that the error in a causal model is unrelated to $X$.
Name the groups and shares.
$P(X=1) = 0.4, \ P(X=0) = 0.6$
Shares add to one.
Read the conditional means.
$E[Y \mid 1] = 60, \ E[Y \mid 0] = 45$
Mean wage in each group.
Weight the first mean.
$0.4 \times 60 = 24$
Its share of the total.
Weight the second mean.
$0.6 \times 45 = 27$
Its share of the total.
Add the pieces.
$24 + 27 = 51$
Between $45$ and $60$, as it must be.
Define the dummy.
$X = 1 \text{ if union}$
Two values only.
Write the regression.
$E[Y \mid X] = \beta_0 + \beta_1 X$
Two unknowns.
Read the non-member mean.
$E[Y \mid 0] = 36$
Dollars an hour.
Match the intercept.
$\beta_0 = 36$
The line at $X = 0$.
Read the member mean.
$E[Y \mid 1] = 44$
Dollars an hour.
Solve for the slope.
$\beta_1 = 44 - 36 = 8$
The difference in means.
List the CEF points.
$(10, 15), (12, 18), (14, 21)$
Years and mean wage.
Compute the first slope.
$\tfrac{18 - 15}{12 - 10} = 1.5$
Rise over run.
Compute the second slope.
$\tfrac{21 - 18}{14 - 12} = 1.5$
The same slope again.
Conclude the CEF is linear.
$E[Y \mid X] = \beta_0 + 1.5X$
Regression equals the CEF.
Solve for the intercept.
$15 - 1.5 \times 10 = 0$
The line passes through each mean.
Write the line.
$E[Y \mid X] = 1.5X$
Dollars an hour.
Interpret without overclaiming.
$\text{descriptive slope}$
Not yet a causal return to schooling.
Weight the exporters.
$0.2 \times 75 = 15$
Their share of the mean.
Weight the others.
$0.8 \times 55 = 44$
Their share of the mean.
Add the pieces.
In a town, $40\%$ of workers belong to a union and earn a mean of $60$ dollars an hour; the other $60\%$ earn a mean of $45$. What is the mean hourly wage of all workers?
Answer: dollars an hour
Complete the worked solution: $60\%$ of households in a county own a home. Owners spend a mean of $41$ hundred dollars a month on housing and renters $37$. Find each weighted piece and the county mean.
Weight the owners' mean.
$0.6 \times 41 =$ a
Owners' share of the county mean.
Weight the renters' mean.
$0.4 \times 37 =$ b
Renters' share of the county mean.
Add the two pieces.
$E[Y] =$ c
The law of iterated expectations.
A regression of hourly wage on a union dummy ($X = 1$ for members) uses a large sample in which members average $63$ dollars and non-members $56$. What is the estimated slope on the union dummy?
Answer: dollars an hour
Match each object to its definition.
| the mean of Y at each value of X | Y minus its mean given X; averages zero at every X | the straight line closest to the CEF in mean squared error | the mean of Y over everyone, ignoring X | |
|---|---|---|---|---|
| conditional expectation function | ||||
| CEF error | ||||
| population regression line | ||||
| unconditional mean |
Four workers have ($X$ = years of experience, $Y$ = wage): $(0, 19)$, $(0, 27)$, $(5, 28)$ and $(5, 36)$. Fill in the sample CEF at each $X$ and the overall mean.
| value | |
|---|---|
| mean wage at X = 0 | |
| mean wage at X = 5 | |
| overall mean wage |
In a large survey, mean hourly wages by years of schooling are $14$ dollars at $8$ years, $20$ at $12$ years and $26$ at $16$ years. These are the only schooling levels in the population. What slope does the population regression of wages on years of schooling have, in dollars per year?
Answer: dollars per year
A labor economist regresses weekly pay (in tens of dollars) on a dummy $F = 1$ for women, using a large national sample. Women's mean pay is $51$ and men's is $60$. What is the coefficient on $F$?
Answer: tens of dollars
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Four workers have ($X$ = years of experience, $Y$ = wage): $(0, 12)$, $(0, 22)$, $(5, 19)$ and $(5, 29)$. Fill in the sample CEF at each $X$ and the overall mean.
| value | |
|---|---|
| mean wage at X = 0 | |
| mean wage at X = 5 | |
| overall mean wage |
You can connect regression to conditional means. Explain to someone why a regression on a union dummy computes the same gap as comparing two averages.
16. Your turn: 20% of firms export, with mean profit 75; the rest have mean profit 55., step 3
$15 + 44 = 59$
The mean profit of all firms.