Back to the on-screen lesson ·

The conditional expectation function

Conditional means, the law of iterated expectations, and regression as the best linear approximation to the CEF; a dummy's slope is a difference in means.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute conditional and overall means, read a dummy-variable regression as a difference in means, and find the regression line when the CEF is linear.

2. What you already have

You can compute a mean and a weighted average, and from the last lesson you know that a comparison of group means answers a causal question only when selection bias is zero. In statistics you fitted a least-squares line to a scatterplot.

This lesson says precisely what that line is trying to estimate, before any question of causation. The object is the conditional expectation function, and regression is the best straight-line summary of it.

3. Terms to use precisely

TermWhat it means
Conditional expectation$E[Y \mid X = x]$: the population mean of $Y$ among units with $X = x$.
CEFThe conditional expectation function: the list of conditional means, one for each $x$.
CEF error$e = Y - E[Y \mid X]$; its mean is zero at every value of $X$.
Law of iterated expectations$E[Y] = E\big[E[Y \mid X]\big]$: the overall mean is the weighted average of the conditional means.
Population regressionThe line $\beta_0 + \beta_1 X$ that minimizes $E[(Y - \beta_0 - \beta_1 X)^2]$.
Dummy variableA regressor that equals 1 for one group and 0 for the other.
RegressorThe variable $X$ on the right-hand side; $Y$ is the dependent variable.

4. Regression approximates the conditional mean

For every value $x$ of a variable $X$, the units with that value have some average outcome, $E[Y \mid X = x]$. Taken together, these averages form the conditional expectation function (CEF). It is the best possible prediction of $Y$ from $X$: no other function of $X$ has a smaller mean squared prediction error.

Any outcome can then be written as its conditional mean plus an error:

$$Y = E[Y \mid X] + e, \qquad E[e \mid X] = 0.$$

That split is not an assumption. It holds by construction, because $e$ is defined as whatever is left over. The error averages zero at every $X$ simply because the mean was subtracted there.

The population regression line $\beta_0 + \beta_1 X$ is the straight line closest to the CEF. When the CEF is itself a straight line, regression recovers it exactly. When it curves, regression gives its best linear approximation. When $X$ is a dummy, the CEF has only two points, a line always fits them exactly, and the slope is simply the difference in the two group means.

Another way: action

Sort a class by hours of sleep, compute the average test score at each number of hours, and plot those averages. The dots are the sample CEF; a regression line through the raw data runs close to them.

Another way: steps

  1. Group the units by their value of $X$.
  2. Average $Y$ within each group: that is the CEF at that $x$.
  3. Weight the group means by group shares to recover the overall mean.
  4. Check whether the conditional means lie on a line.
  5. Read the regression slope as the change in the mean per unit of $X$ (exactly so if the CEF is linear, approximately if not).

5. Why the law of iterated expectations holds

Suppose $30\%$ of workers are in a union and earn $50$ dollars an hour on average, and $70\%$ are not and earn $40$. The overall mean is not the simple average $45$, because the groups are different sizes. Think of $100$ workers: $30$ of them earn a total of $1{,}500$ and $70$ earn $2{,}800$, so the total is $4{,}300$ and the mean is $43$.

That calculation is $0.3 \times 50 + 0.7 \times 40 = 43$, which is the law of iterated expectations: $E[Y] = \sum_x P(X = x)\,E[Y \mid X = x]$. Averaging the conditional means with the right weights must give back the overall mean, because every unit is counted exactly once. The same law is why the CEF error has mean zero overall: it has mean zero within every group, and a weighted average of zeros is zero.

6. Why a dummy regression is a difference in means

Let $X$ be $1$ for union members and $0$ otherwise. The regression line $\beta_0 + \beta_1 X$ takes only two values: $\beta_0$ at $X = 0$ and $\beta_0 + \beta_1$ at $X = 1$. The line that best fits two conditional means passes through both of them, because two points always lie on a line.

So $\beta_0 = E[Y \mid X = 0]$ and $\beta_0 + \beta_1 = E[Y \mid X = 1]$, which gives $\beta_1 = E[Y \mid X = 1] - E[Y \mid X = 0]$. The regression slope is exactly the naive difference from the last lesson. That is the first sign of a theme that runs through the course: a regression coefficient is only as causal as the comparison it makes. Running a regression does not remove selection bias; with a dummy regressor it computes the same naive gap.

7. When the CEF is not a line

Wages rise quickly with the first years of experience and then flatten; test scores may rise with sleep up to eight hours and then level off. In cases like these the CEF curves, and a straight line cannot match every conditional mean.

Regression still gives the line that is closest to the CEF on average, weighting each value of $X$ by how common it is. Its slope is then a kind of average slope of the CEF. That is often useful, but it can hide important structure: a slope of zero can come from a CEF that rises and then falls. Later lessons show how to let the line bend, with logs, squares and interactions, when the shape matters.

A practical habit follows. Before trusting a single slope, look at the conditional means themselves: bin the regressor into groups, average the outcome in each, and plot the averages. If the dots fall close to a line, one slope is a fair summary. If they bend, rise and fall, or jump at some value, the slope is hiding something the reader of your results deserves to see, and a more flexible specification is worth the extra coefficients.

8. Why the CEF is the best predictor

Suppose you must guess a worker's wage knowing only their years of schooling, and you are penalized by the square of your error. Whatever you guess for workers with twelve years, the guess that makes the average squared error smallest among them is their mean wage: any other number adds the squared distance between it and the mean to every penalty.

Doing that separately at every value of schooling gives exactly the CEF. So no function of $X$, however complicated, predicts $Y$ better in mean squared error. This is why regression, which approximates the CEF with a line, is the natural starting point for prediction as well as for description.

It also shows what the CEF cannot do. It is the best guess given the $X$ you observe. It says nothing about what would happen to $Y$ if you reached in and changed $X$ for a particular person, which is the causal question from the last lesson. The same numbers answer the prediction question well and the causal one only under extra assumptions.

9. Working a CEF problem, step by step

A survey records mean hourly wages of $20$ dollars at $12$ years of schooling, $24$ at $14$ years and $28$ at $16$ years, and no other schooling levels.

  1. List the CEF. Three points: $(12, 20)$, $(14, 24)$, $(16, 28)$.
  2. Check whether they lie on a line. From $12$ to $14$ the mean rises $4$ over $2$ years, a slope of $2$; from $14$ to $16$ it rises $4$ over $2$, also $2$.
  3. Apply the linear-CEF fact. The regression line equals the CEF, so $\beta_1 = 2$ dollars per year.
  4. Find the intercept. $20 - 2 \times 12 = -4$, so the line is $-4 + 2X$.
  5. Interpret carefully. Mean wages are two dollars higher per extra year of schooling in this population. Whether a year of schooling causes two more dollars is a separate question about selection.

10. How to check a CEF answer

Three checks catch most errors.

  1. Is an overall mean between the group means? A weighted average can never fall outside the range of what it averages. If it does, a weight is wrong.
  2. Do the weights add to one? Shares of $30\%$ and $70\%$ do; if yours add to more, a group has been counted twice.
  3. Does the dummy slope have the right sign? It is the $X = 1$ mean minus the $X = 0$ mean. If the members earn more, the slope must be positive.

And a fourth, conceptual check: have you described the answer as a difference in means, not as an effect? The CEF describes the population; causation needs an argument about selection.

11. In the world: the gender pay gap

The Bureau of Labor Statistics reports that women working full time in the United States earn roughly $83$ cents for every dollar men earn, measured by median weekly earnings. A regression of pay on a dummy for being a woman reproduces exactly that kind of gap: the coefficient is the women's mean minus the men's mean.

That number is a CEF fact about the population, and it is useful: it tells you how different the average paychecks are. It is not by itself the effect of discrimination, because men and women in the data also differ, on average, in hours, occupations and years of experience. Some of those differences may themselves reflect discrimination earlier in life, which is why economists argue about which to hold fixed.

Later lessons add those variables to the regression and show how the coefficient changes. The point of this lesson is the starting place: a dummy coefficient is a difference in conditional means, and interpreting it causally needs an argument, not a regression command.

In the units of the transfer problem: if women's mean weekly pay is $100$ tens of dollars and men's is $120$, the coefficient on the dummy is $-20$, a gap of two hundred dollars a week, or about seventeen percent of men's pay.

12. A regression line is not automatically an effect

The most common mistake is reading a regression slope as what would happen if $X$ were changed. The slope describes how the conditional mean of $Y$ varies with $X$ in the population; it is causal only when the comparison it makes is free of selection.

A second mistake is averaging group means without weights. The overall mean weights each group by its share; the simple average of two group means is right only when the groups are the same size.

A third is treating $E[e \mid X] = 0$ as an assumption about the world. For the CEF error it holds by definition. What is an assumption, later, is that the error in a causal model is unrelated to $X$.

13. The overall mean from two groups

  1. Name the groups and shares.

    $P(X=1) = 0.4, \ P(X=0) = 0.6$

    Shares add to one.

  2. Read the conditional means.

    $E[Y \mid 1] = 60, \ E[Y \mid 0] = 45$

    Mean wage in each group.

  3. Weight the first mean.

    $0.4 \times 60 = 24$

    Its share of the total.

  4. Weight the second mean.

    $0.6 \times 45 = 27$

    Its share of the total.

  5. Add the pieces.

    $24 + 27 = 51$

    Between $45$ and $60$, as it must be.

14. A dummy regression

  1. Define the dummy.

    $X = 1 \text{ if union}$

    Two values only.

  2. Write the regression.

    $E[Y \mid X] = \beta_0 + \beta_1 X$

    Two unknowns.

  3. Read the non-member mean.

    $E[Y \mid 0] = 36$

    Dollars an hour.

  4. Match the intercept.

    $\beta_0 = 36$

    The line at $X = 0$.

  5. Read the member mean.

    $E[Y \mid 1] = 44$

    Dollars an hour.

  6. Solve for the slope.

    $\beta_1 = 44 - 36 = 8$

    The difference in means.

15. A linear CEF in schooling

  1. List the CEF points.

    $(10, 15), (12, 18), (14, 21)$

    Years and mean wage.

  2. Compute the first slope.

    $\tfrac{18 - 15}{12 - 10} = 1.5$

    Rise over run.

  3. Compute the second slope.

    $\tfrac{21 - 18}{14 - 12} = 1.5$

    The same slope again.

  4. Conclude the CEF is linear.

    $E[Y \mid X] = \beta_0 + 1.5X$

    Regression equals the CEF.

  5. Solve for the intercept.

    $15 - 1.5 \times 10 = 0$

    The line passes through each mean.

  6. Write the line.

    $E[Y \mid X] = 1.5X$

    Dollars an hour.

  7. Interpret without overclaiming.

    $\text{descriptive slope}$

    Not yet a causal return to schooling.

16. Your turn: 20% of firms export, with mean profit 75; the rest have mean profit 55.

  1. Weight the exporters.

    $0.2 \times 75 = 15$

    Their share of the mean.

  2. Weight the others.

    $0.8 \times 55 = 44$

    Their share of the mean.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Add the pieces.

17. Guided practice

In a town, $40\%$ of workers belong to a union and earn a mean of $60$ dollars an hour; the other $60\%$ earn a mean of $45$. What is the mean hourly wage of all workers?

Answer: dollars an hour

18. Guided practice

Complete the worked solution: $60\%$ of households in a county own a home. Owners spend a mean of $41$ hundred dollars a month on housing and renters $37$. Find each weighted piece and the county mean.

  1. Weight the owners' mean.

    $0.6 \times 41 =$ a

    Owners' share of the county mean.

  2. Weight the renters' mean.

    $0.4 \times 37 =$ b

    Renters' share of the county mean.

  3. Add the two pieces.

    $E[Y] =$ c

    The law of iterated expectations.

19. Guided practice

A regression of hourly wage on a union dummy ($X = 1$ for members) uses a large sample in which members average $63$ dollars and non-members $56$. What is the estimated slope on the union dummy?

Answer: dollars an hour

20. Practice

Match each object to its definition.

the mean of Y at each value of XY minus its mean given X; averages zero at every Xthe straight line closest to the CEF in mean squared errorthe mean of Y over everyone, ignoring X
conditional expectation function
CEF error
population regression line
unconditional mean

21. Practice

Four workers have ($X$ = years of experience, $Y$ = wage): $(0, 19)$, $(0, 27)$, $(5, 28)$ and $(5, 36)$. Fill in the sample CEF at each $X$ and the overall mean.

value
mean wage at X = 0
mean wage at X = 5
overall mean wage

22. Practice

In a large survey, mean hourly wages by years of schooling are $14$ dollars at $8$ years, $20$ at $12$ years and $26$ at $16$ years. These are the only schooling levels in the population. What slope does the population regression of wages on years of schooling have, in dollars per year?

Answer: dollars per year

23. Somewhere new

A labor economist regresses weekly pay (in tens of dollars) on a dummy $F = 1$ for women, using a large national sample. Women's mean pay is $51$ and men's is $60$. What is the coefficient on $F$?

Answer: tens of dollars

24. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

25. Test question

Four workers have ($X$ = years of experience, $Y$ = wage): $(0, 12)$, $(0, 22)$, $(5, 19)$ and $(5, 29)$. Fill in the sample CEF at each $X$ and the overall mean.

value
mean wage at X = 0
mean wage at X = 5
overall mean wage

26. What you can do now

You can connect regression to conditional means. Explain to someone why a regression on a union dummy computes the same gap as comparing two averages.

Working for the steps left to you

16. Your turn: 20% of firms export, with mean profit 75; the rest have mean profit 55., step 3

$15 + 44 = 59$

The mean profit of all firms.