Back to the on-screen lesson ·

Heteroskedasticity and robust standard errors

When error spread varies with x, OLS stays unbiased but its usual standard errors do not: compute robust errors, test with LM = nR², and weight when the pattern is known.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute a robust standard error, run the LM test for heteroskedasticity, and transform data for weighted least squares.

2. What you already have

From the standard-error lesson you know the formula $se = \hat\sigma/\sqrt{S_{xx}}$ and the fifth assumption behind it, homoskedasticity: the errors have the same variance at every value of the regressors. You also know that unbiasedness needed only the first four assumptions.

In economic data the fifth assumption fails constantly. Spending varies more among rich households than poor ones; wages vary more among graduates than among dropouts; large firms' profits swing by more dollars than small firms'. This lesson shows what goes wrong, how to detect it and how to fix the inference, and it ends with the practical rule most economists follow.

3. Terms to use precisely

TermWhat it means
HeteroskedasticityError variance that changes with the regressors: $Var(u \mid x)$ depends on $x$.
HomoskedasticityError variance that is the same at every value of the regressors.
Robust standard errorA standard error valid whether or not the errors are homoskedastic (White or Huber-White).
Breusch-Pagan testA test that regresses squared residuals on the regressors; $LM = nR^2$.
LM statisticA Lagrange multiplier statistic, compared with a chi-squared critical value.
Weighted least squaresOLS on data divided by each observation's error standard deviation.
EfficiencyHaving the smallest variance among a class of unbiased estimators.

4. Keep the estimate, fix the standard error

Heteroskedasticity does not bias OLS: the slope still equals the true slope plus a weighted sum of errors with mean zero. What changes is the variance of that sum. With error variance $\sigma_i^2$ for observation $i$, the slope's true variance is

$$Var(\hat\beta_1) = \frac{\sum (x_i - \bar x)^2 \sigma_i^2}{S_{xx}^2},$$

which reduces to $\sigma^2/S_{xx}$ only when every $\sigma_i^2$ is the same. The robust (White) standard error estimates each $\sigma_i^2$ by the observation's own squared residual:

$$se_{robust}(\hat\beta_1) = \frac{\sqrt{\sum (x_i - \bar x)^2 \hat u_i^2}}{S_{xx}}.$$

It is valid in large samples whether the errors are homoskedastic or not. Using it changes the t statistics and confidence intervals, never the coefficients.

To test for heteroskedasticity, regress the squared residuals on the regressors. If the spread does not depend on them, this regression explains nothing. The Breusch-Pagan statistic is $LM = n \cdot R^2$, compared with a chi-squared critical value with as many degrees of freedom as regressors.

Another way: action

Plot the residuals against $x$. If they form a band of constant width, the errors look homoskedastic. If the band fans out or narrows, they do not.

Another way: steps

  1. Fit OLS and compute the residuals.
  2. Weight each squared residual by its squared deviation $(x_i - \bar x)^2$.
  3. Sum the weighted squares and take the square root.
  4. Divide by $S_{xx}$ for the robust standard error.
  5. Use it in every t statistic and confidence interval.

5. Why the usual formula goes wrong

The usual formula averages all the squared residuals into one $\hat\sigma^2$ and treats every observation as equally noisy. The true variance of the slope, though, weights each observation's noise by $(x_i - \bar x)^2$: points far from the mean of $x$ have the most leverage on the slope, so their noise matters most.

If the noisiest observations are exactly those far from the mean — the fan shape typical of income and wage data — the usual formula spreads their large variance across all points and understates the slope's uncertainty. The t statistics are too large and the confidence intervals too narrow, and the tests reject true nulls more often than the stated 5 percent. If the noisiest points are near the middle, the usual formula overstates the uncertainty instead. Either way, the stated level of the test is wrong.

6. Why the robust formula works

Any single squared residual is a poor estimate of its own observation's error variance: it is one number, and it could be large or small by chance. But the robust formula never needs any single estimate to be good. It needs only the weighted sum $\sum (x_i - \bar x)^2 \hat u_i^2$ to be close to $\sum (x_i - \bar x)^2 \sigma_i^2$, and in a large sample the chance errors in the individual terms average out.

That is why robust standard errors are a large-sample tool. With a few dozen observations they can be unreliable, usually too small; with hundreds or thousands they are dependable. Software offers small-sample corrections, often labeled HC1 to HC3, that scale the squared residuals up slightly.

7. Why test at all?

If robust standard errors are valid either way, why test for heteroskedasticity? Most applied economists no longer do before computing standard errors: they report robust errors by default, as a matter of course.

Testing still has two uses. First, the pattern of heteroskedasticity is information: spending that varies more among the rich tells you something about consumption behavior. Second, if the variance pattern is known, weighted least squares gives a more precise estimate than OLS, and the test helps decide whether that is worth doing.

The Breusch-Pagan test is simple. Under homoskedasticity the squared residuals are unrelated to the regressors, so regressing them on the regressors gives $R^2$ near zero, and $n R^2$ has a chi-squared distribution. A large $n R^2$ — above about $3.84$ with one regressor, $5.99$ with two — rejects homoskedasticity at 5 percent. A related version, the White test, adds the squares and cross-products of the regressors to the auxiliary regression, so it can detect variance that changes in curved ways as well as straight ones, at the cost of using up more degrees of freedom in the auxiliary regression.

8. Weighted least squares

If you know how the variance changes — say $Var(u \mid x) = \sigma^2 x$, as it might for spending and income — you can remove the heteroskedasticity. Divide every term of the equation, including the intercept's column of ones, by $\sqrt{x}$. The new error $u/\sqrt{x}$ has variance $\sigma^2$ everywhere, and OLS on the transformed data is again the best linear unbiased estimator.

The coefficients keep their meaning; only the variables are rescaled. The gain is precision: WLS downweights the noisy observations, which carry less information about the slope. The risk is getting the variance pattern wrong, in which case WLS can be less precise than OLS. That is why many researchers use WLS only when the pattern is clear from how the data were built — averages of groups of different sizes, for instance — and use OLS with robust errors otherwise.

9. Beyond heteroskedasticity: clustered errors

Robust standard errors allow each observation its own error variance, but they still assume the errors of different observations are unrelated. Often they are not. Students in the same classroom share a teacher; workers in the same state face the same labor market; repeated observations of one firm share its management. Their errors move together.

When errors are correlated within groups, both the usual and the robust standard errors understate uncertainty, sometimes badly. Ten thousand students in fifty schools carry much less independent information than ten thousand students drawn from ten thousand schools, especially when the treatment itself is assigned by school. Clustered standard errors extend the robust formula by summing residuals within each group before squaring, so shared shocks are counted once per group rather than once per person.

The practical rule is to cluster at the level at which the treatment varies or the shocks are shared — the school, the state, the firm — and to be wary when there are only a handful of clusters, since the formula, like the robust one, is a large-sample tool whose "sample size" is now the number of clusters.

10. Working a robust standard error, step by step

Four observations have $x = 1, 2, 4, 5$ and $y = 11, 16, 16, 21$; OLS gives $\hat y = 10 + 2x$.

  1. Deviations. $\bar x = 3$, so $x - \bar x = -2, -1, 1, 2$ and $S_{xx} = 10$.
  2. Residuals. Fitted values $12, 14, 18, 20$; residuals $-1, 2, -2, 1$.
  3. Weighted sum. $4(1) + 1(4) + 1(4) + 4(1) = 16$.
  4. Root. $\sqrt{16} = 4$.
  5. Robust SE. $4/10 = 0.4$.

The usual formula gives $\hat\sigma^2 = 10/2 = 5$ and $se = \sqrt{5/10} \approx 0.71$. Here the biggest residuals sit near the middle, so the robust error is smaller than the usual one.

11. How to check a heteroskedasticity answer

Three checks.

  1. Did the coefficient change? It must not. Only standard errors, t statistics and intervals change with robust errors.
  2. Did you weight by squared deviations? Without the weights, you have computed something like the usual formula.
  3. Did you use $n$, not $n - k - 1$, in the LM statistic? It is the sample size times the auxiliary $R^2$.

And a reporting check: say which standard errors a table uses. A reader cannot otherwise tell whether a t statistic of $2.5$ is trustworthy.

12. In the world: spending, income and the fan-shaped cloud

The Consumer Expenditure Survey records what thousands of US households spend. Plot annual spending on food away from home against income and the cloud fans out: households earning $30{,}000$ dollars spend within a narrow range, while those earning $200{,}000$ range from very little to a great deal. The error variance grows with income.

A regression of spending on income still estimates the average slope without bias, perhaps $2$ cents per extra dollar of income. But its usual standard error might be $0.002$, giving $t = 10$, while the robust standard error is $0.004$, giving $t = 5$. Both reject zero, but the robust interval is twice as wide, and a claim that the slope is "between $1.6$ and $2.4$ cents" would be overconfident.

In a study of a smaller effect the difference decides the result. A slope with a usual t of $2.4$ and a robust t of $1.5$ is significant or not depending on which standard error is used, and only the robust one is valid here. That is why the standard practice in economics journals is to report robust standard errors for cross-section regressions unless there is a specific reason not to, and to say so in every table note. When observations are grouped — households in the same county, say — the same journals expect standard errors clustered by group, for the reasons given above, and readers should look for which kind a table reports before trusting its stars.

13. Heteroskedasticity does not bias the coefficients

The most common mistake is to believe heteroskedasticity makes OLS estimates wrong. The estimates are still unbiased; only the usual standard errors and tests are invalid.

A second mistake is thinking robust standard errors are always larger. They are larger when the noisiest observations are far from the mean of $x$, and smaller when they are near it.

A third is treating robust standard errors as a fix for bias. They correct inference under heteroskedasticity only; they do nothing about omitted variables or selection.

14. Robust SE from a summary sum

  1. Read the weighted sum.

    $\sum (x_i - \bar x)^2 \hat u_i^2 = 144$

    From the residuals.

  2. Read the spread.

    $S_{xx} = 40$

    Squared deviations of $x$.

  3. Take the square root.

    $\sqrt{144} = 12$

    The numerator.

  4. Divide by the spread.

    $12 \div 40 = 0.3$

    Robust standard error.

  5. Use it for inference.

    $t = \hat\beta \div 0.3$

    In place of the usual error.

15. The Breusch-Pagan test

  1. Fit the main regression.

    $\hat u_i$

    Keep the residuals.

  2. Square the residuals.

    $\hat u_i^2$

    Estimates of each variance.

  3. Regress them on x.

    $R^2_{aux} = 0.04$

    Does the spread depend on $x$?

  4. Read the sample size.

    $n = 300$

    All observations.

  5. Multiply the two.

    $LM = 300 \times 0.04 = 12$

    The test statistic.

  6. Compare with chi-squared.

    $12 > 3.84$

    Reject homoskedasticity.

16. Weighted least squares for a known pattern

  1. State the variance pattern.

    $Var(u \mid x) = \sigma^2 x$

    Spreads with income.

  2. Find the error's scale.

    $\sqrt{x}$

    Its standard deviation factor.

  3. Divide the equation.

    $\dfrac{y}{\sqrt{x}} = \dfrac{\beta_0}{\sqrt{x}} + \beta_1\sqrt{x} + \dfrac{u}{\sqrt{x}}$

    Every term.

  4. Check the new error.

    $Var(u/\sqrt{x}) = \sigma^2$

    Constant now.

  5. Transform one household.

    $x = 16, \ y = 40 \Rightarrow 40 \div 4 = 10$

    Divide by 4.

  6. Run OLS on the new data.

    $\text{no intercept column; } 1/\sqrt{x} \text{ instead}$

    Same coefficients, new regressors.

  7. Read the result.

    $\hat\beta_1 \text{ keeps its meaning}$

    More precise if the pattern is right.

17. Your turn: weighted sum 900, S_xx = 60.

  1. Take the square root.

    $\sqrt{900} = 30$

    The numerator.

  2. Divide by the spread.

    $30 \div 60 = 0.5$

    Robust standard error.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Use it in the t statistic.

18. Guided practice

To test for heteroskedasticity, an analyst regresses the squared OLS residuals on the regressors, using all $200$ observations. This auxiliary regression has $R^2 = 0.05$. What is the LM statistic?

Answer: LM statistic

19. Guided practice

Complete the worked solution: household spending $y$ has error variance $\sigma^2 x$, where $x$ is income. One household has $x = 81$ and $y = 324$, and a coefficient on income of $8$ is being estimated. Find the weight divisor, the transformed $y$, and the transformed income term.

  1. Find the error's scale.

    $\sqrt{x} = \sqrt{81} =$ a

    Dividing by it makes the error variance constant.

  2. Transform the outcome.

    $y / \sqrt{x} =$ b

    Every term is divided by the same number.

  3. Transform the income term.

    $8 \times x / \sqrt{x} =$ c

    The slope is unchanged; the regressor becomes $\sqrt{x}$.

20. Guided practice

In a simple regression, $\sum (x_i - \bar x)^2 \hat u_i^2 = 36$ and $S_{xx} = 20$. What is the heteroskedasticity-robust standard error of the slope?

Answer: robust standard error

21. Practice

For $x = 1, 2, 4, 5$ and $y = 11, 16, 16, 21$ with OLS line $\hat y = 10 + 2x$, fill in the weighted sum of squared residuals, its square root and the robust standard error of the slope.

value
weighted sum of squared residuals
its square root
robust standard error

22. Practice

The errors in a wage regression are heteroskedastic. Match each claim to a verdict.

truefalse
The OLS slope is still unbiased.
The usual t statistics are still valid.
Robust standard errors give valid large-sample tests.
Robust standard errors change the slope estimate.

23. Practice

Four observations have $x = 1, 2, 4, 5$ and $y = 19, 30, 22, 33$. The OLS line is $\hat y = 20 + 2x$. What is the heteroskedasticity-robust standard error of the slope?

Answer: robust standard error

24. Somewhere new

A study of housing prices across US metro areas reports a slope of $2$ on median income. Its usual standard error is $0.2$, but the residuals fan out with income, so the authors also compute a robust standard error of $0.4$. What is the t statistic using the robust standard error?

Answer: robust t statistic

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

For $x = 1, 2, 4, 5$ and $y = 11, 16, 16, 21$ with OLS line $\hat y = 10 + 2x$, fill in the weighted sum of squared residuals, its square root and the robust standard error of the slope.

value
weighted sum of squared residuals
its square root
robust standard error

27. What you can do now

You can handle unequal error spreads. Explain to someone why heteroskedasticity leaves a slope unbiased but can make its t statistic misleading.

Working for the steps left to you

17. Your turn: weighted sum 900, S_xx = 60., step 3

$t = \hat\beta \div 0.5$

Valid under heteroskedasticity.