Back to the on-screen lesson ·
Read coefficients as partial effects, compute predicted differences, see the partialling-out result, and penalize fit with adjusted R-squared.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to predict and compare outcomes with several regressors, compute a coefficient by partialling out, and use adjusted R-squared.
You can fit, read and test a simple regression, and you know its slope is biased whenever something else that affects $y$ moves with $x$. In the class-size data, richer districts have smaller classes and higher scores, so the simple slope mixes the effect of class size with the effect of income.
Multiple regression is the first tool for that problem. It puts income in the regression alongside class size, so the class-size coefficient compares districts with the same income. This lesson shows what that means, how the coefficient is computed, and how fit is measured when there are many regressors.
| Term | What it means |
|---|---|
| Multiple regression | $y = \beta_0 + \beta_1 x_1 + \dots + \beta_k x_k + u$, with $k$ regressors. |
| Partial effect | $\beta_j$: the change in $E[y]$ for a unit change in $x_j$, other regressors held fixed. |
| Control variable | A regressor included so the coefficient of interest compares like with like. |
| Partialling out | Getting $\hat\beta_1$ by regressing $y$ on the part of $x_1$ not predicted by the other regressors. |
| Degrees of freedom | $n - k - 1$: observations minus estimated coefficients. |
| Adjusted R-squared | $1 - (1 - R^2)(n - 1)/(n - k - 1)$: fit penalized for the number of regressors. |
| Ceteris paribus | Latin for "other things equal"; what a partial effect holds. |
With two regressors the fitted model is $\hat y = \hat\beta_0 + \hat\beta_1 x_1 + \hat\beta_2 x_2$. OLS still minimizes the sum of squared residuals, now over three coefficients, and the solution has a clean interpretation. The predicted change when both regressors change is
$$\Delta \hat y = \hat\beta_1 \Delta x_1 + \hat\beta_2 \Delta x_2.$$
Set $\Delta x_2 = 0$ and $\hat\beta_1$ is the predicted change in $y$ per unit of $x_1$ holding $x_2$ fixed. That is what a control does: the class-size coefficient now compares districts that differ in class size but not in income.
The partialling-out result (the Frisch–Waugh–Lovell theorem) makes this exact. Regress $x_1$ on $x_2$ and keep the residuals $\tilde r_1$ — the part of class size that income does not predict. Then $\hat\beta_1$ is the simple slope of $y$ on $\tilde r_1$:
$$\hat\beta_1 = \frac{\sum \tilde r_{1i} y_i}{\sum \tilde r_{1i}^2}.$$
The coefficient uses only the variation in $x_1$ that is unrelated to $x_2$.
Another way: action
Sort districts into income bands. Within one band, plot scores against class size and draw a line. Do this in every band and average the slopes: that is roughly what the class-size coefficient in a regression with income does.
Another way: steps
The normal equations say the residuals are uncorrelated with every regressor. Take the part of $x_1$ that $x_2$ predicts: it is a linear function of $x_2$, so it is uncorrelated with the residuals and with $\tilde r_1$. Only the leftover part $\tilde r_1$ can carry information about $\beta_1$ that is not already carried by $x_2$.
The result has two practical lessons. First, a control only helps with variation it can predict: if class size varied only because of income, $\tilde r_1$ would be zero everywhere and $\beta_1$ could not be estimated at all. That is perfect collinearity, and software drops a variable when it occurs. Second, when $x_1$ and $x_2$ are highly but not perfectly correlated, $\tilde r_1$ has little variation left, and the standard error of $\hat\beta_1$ is large. That is multicollinearity: not a bias, but a loss of precision.
Adding income as a control removes the part of the class-size slope that runs through income. It does nothing about other differences between districts that the model leaves out — parental education, teacher quality, the share of students learning English. If those still move with class size once income is held fixed, the coefficient is still biased.
So the zero conditional mean assumption becomes $E[u \mid x_1, x_2] = 0$: nothing left in the error may move with any regressor. Multiple regression turns the causal question into a question about which variables have been controlled for, and whether any important one is missing. The next lesson measures exactly how much a missing variable biases the slope.
Controls can also hurt. A variable that is itself caused by the treatment — controlling for occupation when estimating the return to schooling, for example — removes part of the effect you want to measure. Good controls are determined before the treatment.
The partialling-out result also explains how a control changes a standard error. The simple-regression formula was $\hat\sigma/\sqrt{S_{xx}}$. In a multiple regression, $S_{xx}$ is replaced by the sum of squared residuals $\tilde r_1$, which equals $S_{xx}(1 - R_1^2)$, where $R_1^2$ is the $R^2$ from regressing $x_1$ on the other regressors.
So a control that is strongly related to $x_1$ throws away most of the variation the slope was estimated from. If income explains $90$ percent of the variation in class size, only a tenth of $S_{xx}$ is left, and the standard error grows by a factor of $\sqrt{10}$, about three. At the same time the control can shrink $\hat\sigma$ by explaining part of $y$. Whether the standard error rises or falls on balance depends on which effect is larger.
This trade-off is worth accepting when the control removes bias: a precise biased number is worse than a noisier unbiased one. It is not worth accepting for variables that neither affect the outcome nor move with the treatment; those only cost degrees of freedom and add noise to every other coefficient in the table.
As the fourth lesson warned, $R^2$ never falls when a regressor is added. Adjusted $R^2$ divides each sum of squares by its degrees of freedom before comparing them:
$$\bar R^2 = 1 - \frac{SSR/(n - k - 1)}{SST/(n - 1)} = 1 - (1 - R^2)\frac{n - 1}{n - k - 1}.$$
Each extra regressor raises the ratio $(n-1)/(n-k-1)$, so it must reduce SSR enough to pay for itself. A useless variable usually lowers $\bar R^2$. With many regressors and few observations the penalty is large, and adjusted $R^2$ can even be negative.
Which variables belong in the regression? The guiding question is what the comparison should hold fixed to answer the causal question, and the answer comes from thinking about how the data were generated, not from which variables happen to be available.
A good control is something that affects the outcome, differs between treated and untreated units, and is determined before the treatment. Family income is a good control for class size: it shapes both, and a district's income does not change because its classes are smaller. Parents' education is another.
A bad control is one that the treatment itself affects. Suppose smaller classes raise attendance, and attendance raises scores. Controlling for attendance would compare districts with the same attendance, removing exactly the channel through which small classes work, and the class-size coefficient would understate the total effect. The same trap appears when economists control for occupation in a wage regression on schooling: schooling changes which occupations people can enter.
Writing down, before running anything, the list of controls and the reason for each is a discipline that prevents the regression from being tuned until it delivers a preferred answer.
The fitted wage equation is $\widehat{\log w} = 1.2 + 0.08\,educ + 0.02\,exper$. Ana has $4$ more years of schooling than Ben and $6$ fewer years of experience.
Using only the schooling coefficient would predict $32$ percent and ignore that Ana spent those years in school rather than at work.
Three checks.
And one check of interpretation: when you say "holding fixed", name the variables held fixed. A coefficient is a partial effect only with respect to the controls actually in the model.
Hedonic price models, used by county assessors and websites like Zillow, regress a home's price on its features. A typical result surprises people: with floor area in the model, the coefficient on the number of bedrooms is often small or negative.
Suppose the model is price $= 40 + 12\,sqft_{100} - 5\,bedrooms$, in thousands of dollars. The $-5$ compares houses with the same floor area: dividing a fixed space into one more bedroom means smaller rooms, and buyers pay a little less for that. It does not say that adding a bedroom lowers a house's value. A new $200$ square foot bedroom changes both regressors, and the model predicts $12 \times 2 - 5 = 19$ thousand dollars more.
Reading coefficients as partial effects is exactly what prevents the wrong renovation advice. Every coefficient in a multiple regression answers a comparison with the other regressors held fixed, and a real decision usually changes several of them at once.
Assessors use the same logic when they value a house that has not sold. They plug its features into the fitted model, and property taxes in many US counties are then set from that predicted price, which is why homeowners who appeal an assessment usually argue about which comparable houses and features the model should have used.
The most common mistake is to believe that adding controls makes a coefficient causal. It removes bias only from the variables actually included; anything omitted that still moves with the regressor keeps biasing it.
A second mistake is reading a coefficient without its "holding fixed" clause. The bedroom coefficient in a house-price model can be negative because it compares houses of the same floor area, where an extra bedroom means smaller rooms.
A third is judging models by $R^2$ when they have different numbers of regressors. Use adjusted $R^2$, or better, a test.
Write the fitted model.
$\hat y = 150 + 4\,sqft_{100} + 6\,baths$
Thousands of dollars.
Convert the floor area.
$2{,}000 \text{ sq ft} = 20$
Hundreds of square feet.
Multiply floor area by its slope.
$4 \times 20 = 80$
Its contribution.
Multiply baths by their slope.
$6 \times 3 = 18$
Three bathrooms.
Add everything to the intercept.
$150 + 80 + 18 = 248$
About 248 thousand dollars.
State the goal.
$\hat\beta_{size} \text{ holding income fixed}$
Class size effect net of income.
Regress class size on income.
$size = \hat a + \hat b\,income + \tilde r$
Keep the residuals.
Read the residual sums.
$\sum \tilde r^2 = 40, \ \sum \tilde r y = -80$
From the data.
Divide the sums.
$-80/40 = -2$
The simple slope on $\tilde r$.
Name the result.
$\hat\beta_{size} = -2$
Points per student, same income.
Compare with the simple slope.
$\text{simple: } -3.5$
Income explained part of it.
Read the sample and model.
$n = 21, \ k = 4$
Four regressors.
Read the R-squared.
$R^2 = 0.6$
Sixty percent explained.
Find the unexplained share.
$1 - 0.6 = 0.4$
What is left.
Count the degrees of freedom.
$n - k - 1 = 16$
And $n - 1 = 20$.
Form the penalty ratio.
$20/16 = 1.25$
The cost of four regressors.
Inflate the unexplained share.
$0.4 \times 1.25 = 0.5$
Penalized.
Subtract from one.
$\bar R^2 = 0.5$
Below $R^2$, as always.
Compute the first part.
$3 \times 4 = 12$
Holding $x_2$ fixed.
Compute the second part.
$-2 \times 5 = -10$
Holding $x_1$ fixed.
Add the two parts.
A fitted model of house prices (thousands of dollars) is $\hat y = 128 + 4\,sqft_{100} + 7\,baths$, with floor area in hundreds of square feet. What is the predicted price of a house with $2100$ square feet and $3$ bathrooms?
Answer: thousand dollars
Complete the worked solution: a regression with $k = 8$ regressors on $n = 49$ observations has $R^2 = 0.85$. Find the unexplained share, the degrees-of-freedom ratio and the adjusted $R^2$.
Find the unexplained share.
$1 - 0.85 =$ a
What the model leaves over.
Form the degrees-of-freedom ratio.
$\dfrac{48}{40} =$ b
$(n - 1)/(n - k - 1)$.
Subtract the penalized share.
$1 - (1 - R^2) \times \text{ratio} =$ c
The adjusted $R^2$.
To get the coefficient on class size in a regression that also includes family income, an analyst first regresses class size on income and keeps the residuals $\tilde r$. She finds $\sum \tilde r^2 = 70$ and $\sum \tilde r \, y = -350$, where $y$ is the test score. What is the class-size coefficient in the multiple regression?
Answer: points per student
A regression with $k = 8$ regressors on $n = 41$ observations has $R^2 = 0.36$. What is its adjusted $R^2$?
Answer: adjusted R-squared
A regression with $2$ regressors and an intercept uses $n = 35$ observations and has $SSR = 192$. Fill in the degrees of freedom, the error variance and the number of estimated coefficients.
| value | |
|---|---|
| coefficients estimated | |
| degrees of freedom | |
| error variance |
In $\widehat{price} = 50 + 12\,sqft_{100} - 8\,bedrooms$, match each comparison to the model's predicted price difference.
| −8 thousand dollars | +12 thousand dollars | +16 thousand dollars | −16 thousand dollars | |
|---|---|---|---|---|
| one more bedroom carved out of the same floor area | ||||
| 100 more square feet, same number of bedrooms | ||||
| a new 200 sq ft bedroom added to the house | ||||
| two more bedrooms carved out of the same floor area |
A county assessor's model is $\widehat{price} = 40 + 9\,sqft_{100} - 5\,bedrooms$ (thousands of dollars). A homeowner plans to build a new bedroom of $400$ square feet onto the house. What change in predicted price does the model imply, in thousands of dollars?
Answer: thousand dollars
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A regression with $3$ regressors and an intercept uses $n = 47$ observations and has $SSR = 387$. Fill in the degrees of freedom, the error variance and the number of estimated coefficients.
| value | |
|---|---|
| coefficients estimated | |
| degrees of freedom | |
| error variance |
You can read a multiple regression. Explain to someone why a bedroom coefficient can be negative even though building a bedroom raises a home's value.
17. Your turn: the fitted model is y = 10 + 3 x1 − 2 x2; x1 rises by 4 and x2 by 5., step 3
$12 - 10 = 2$
The predicted change in $y$.