Back to the on-screen lesson ·
Noise in the outcome costs precision; noise in a regressor shrinks the slope toward zero by the reliability ratio, which can be measured and divided back out.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to compute a reliability ratio, predict an attenuated slope, correct for classical error, and say when the correction does not apply.
From the omitted-variable lesson you know how to sign a bias by asking what is in the error and how it moves with the regressor. From the standard-error lesson you know that noise in the error widens a slope's standard error.
Every economic variable is measured with some error. People misremember their schooling, round their hours, understate their income; firms report inventories with lags; national statistics are revised for years. This lesson asks when such errors matter, which way they bias a regression, and how to correct for them when their size is known.
| Term | What it means |
|---|---|
| Measurement error | The difference between a recorded value and the true value. |
| Classical measurement error | Error with mean zero, unrelated to the true value and to the regression error. |
| Attenuation bias | The pull of a slope toward zero caused by classical error in the regressor. |
| Reliability ratio | $\lambda = \sigma^2_{x^}/(\sigma^2_{x^} + \sigma^2_e)$: the share of a variable's variance that is true signal. |
| Probability limit | The value an estimator approaches as the sample grows, written plim. |
| Validation study | A comparison of reported values with accurate records, used to measure error. |
| Errors in variables | The general name for regressions whose regressors are measured with error. |
Let the true model be $y^ = \beta_0 + \beta_1 x^ + u$.
Error in the outcome. If we observe $y = y^ + v$ with $v$ pure noise, the regression becomes $y = \beta_0 + \beta_1 x^ + (u + v)$. The new error is still unrelated to $x^*$, so OLS is unbiased. The only cost is a larger error variance and wider standard errors.
Error in the regressor. If we observe $x = x^ + e$ with $e$ classical — mean zero and unrelated to $x^$ and $u$ — then $y = \beta_0 + \beta_1 x + (u - \beta_1 e)$. The new error contains $e$, and $x$ contains $e$ too, so the error moves with the regressor. The slope is biased, and in large samples
$$\text{plim } \hat\beta_1 = \lambda\beta_1, \qquad \lambda = \frac{\sigma^2_{x^}}{\sigma^2_{x^} + \sigma^2_e}.$$
Because $0 < \lambda < 1$, the slope keeps its sign and shrinks toward zero: attenuation. The noisier the measure, the smaller $\lambda$ and the stronger the pull.
Another way: action
Take a clean scatterplot with a clear slope and jiggle every point left or right at random. The cloud stretches sideways, and the best-fitting line flattens. That flattening is attenuation.
Another way: steps
The OLS slope is the covariance of $y$ and $x$ divided by the variance of $x$. Classical noise in $x$ does not change the covariance: $e$ is unrelated to $y$, so $Cov(y, x^ + e) = Cov(y, x^) = \beta_1\sigma^2_{x^}$. But it does inflate the variance: $Var(x) = \sigma^2_{x^} + \sigma^2_e$.
Dividing the unchanged covariance by the inflated variance gives $\beta_1\sigma^2_{x^}/(\sigma^2_{x^} + \sigma^2_e) = \lambda\beta_1$. The regression sees extra variation in $x$ that produces no variation in $y$, and concludes $x$ matters less than it does.
Noise in $y$ is different: it adds variation to $y$ that is unrelated to $x$. It leaves both the covariance and the variance of $x$ alone, so the slope is unbiased, just noisier.
The asymmetry is worth remembering because it tells you where to spend effort on data quality. A study can tolerate a rough measure of its outcome if it has a large sample. It cannot tolerate a rough measure of its main regressor, because no sample size undoes attenuation, and the resulting estimate will understate the effect it was designed to measure, however carefully the rest of the study was done. Measure the cause well, above all else, and say how it was measured.
The clean attenuation result needs the error to be pure noise. Real errors often are not. High earners tend to understate income and low earners to overstate it, so the error is negatively related to the true value. People in unhealthy jobs may overstate how much they exercise. Survey respondents round hours to $40$.
When the error is related to the true value or to the outcome, the bias can go in either direction and the simple formula no longer applies. The honest response is to describe how the error arises and reason about its sign, as with omitted variables, or to find better data: administrative records, validation samples, a second independent measurement.
Adding controls usually makes attenuation worse. Recall that a multiple-regression coefficient uses only the part of $x$ not predicted by the other regressors. Controls remove much of the true signal in $x$ — the part related to family background, region, age — but none of the noise, which is unrelated to everything. The remaining variation is a larger share noise, so the effective reliability ratio falls.
Fixed-effects and differencing designs, in later lessons, are especially prone to this. Differences between twins in schooling are small, while each twin's reporting error is as large as ever, so within-twin estimates can be heavily attenuated. That is exactly why twin studies measure the reliability by having each twin report the other's schooling, and correct for it.
If a variable is measured twice, with independent errors, the second measurement can fix the first. Regress $y$ on the first report, but use the second report as an instrument for it — a technique from the instrumental-variables lesson later in the course. Because the second report's error is unrelated to the first's, it carries only the true signal the two share, and the attenuation disappears.
Short of that, the two reports give the reliability directly: their correlation estimates $\lambda$ when both have the same error variance. Dividing the attenuated slope by it gives the corrected estimate. The twins studies use exactly this logic.
When a regressor is measured with classical error and its reliability is unknown, the attenuation result still says something precise: the estimated slope is closer to zero than the true one. So a positive estimate is a lower bound on a positive effect, and the true effect is at least that large.
That can settle an argument. Suppose a study finds that each additional year of schooling raises wages by $6$ percent, using self-reported schooling. A critic argues the true return is small. Measurement error cannot support the critic: it pushes the estimate down, not up. If anything, the true return is larger than $6$ percent. Only other biases, such as omitted ability, could push the other way.
The same logic warns against a common reading of null results. A study that finds no significant relationship between a noisy measure of air pollution and children's test scores has not shown pollution is harmless. With a reliability of, say, $0.4$, a real effect would be shrunk to less than half its size and could easily vanish into the standard error. Before concluding that an effect is absent, ask how well the cause was measured.
True schooling has variance $4$ in a population; self-reports add classical noise of variance $1$. The true return is $0.10$ log points per year.
A fifth of the true return was hidden by noise in the reports.
Three checks.
Then check the assumption: is the error plausibly pure noise? If not, say so, and treat the corrected number as an illustration rather than an estimate.
Finally, check what the reliability ratio was estimated from. A ratio measured in one population, such as a validation study of adults in one state, may not carry over to another sample with different reporting habits. When the correction matters to the conclusion, show the answer for a range of plausible reliabilities rather than for a single value, so the reader can see how much the conclusion leans on that one number.
Identical twins share their genes and upbringing, so comparing twins with different schooling removes much of the ability bias from the omitted-variable lesson. Economists have surveyed thousands of twins at events such as the annual Twins Days festival in Twinsburg, Ohio, and regressed differences in wages on differences in schooling.
The catch is measurement error. Differences in schooling between twins are small — often a year or two — while each twin's reporting error is as large as anyone else's. So a much larger share of the within-twin difference is noise, and the reliability ratio of the difference might be $0.6$ even when each report on its own is $0.9$ reliable.
The researchers solved this by asking each twin to report the other twin's schooling as well. Two independent reports of the same quantity reveal the size of the noise. Uncorrected within-twin estimates of the return were around $7$ percent per year; corrected for measurement error they rose to $10$ percent or more — larger, not smaller, than the simple cross-section estimate. Without the correction, the study would have concluded that ability bias was large; with it, the conclusion reversed. The same data, read with and without attention to noise, told opposite stories about ability.
The most common mistake is to think random measurement error cancels out in a large sample. It cancels out of means, and out of an outcome variable, but in a regressor it biases the slope toward zero however large the sample.
A second mistake is treating a small or insignificant coefficient on a noisy measure as evidence that the true effect is small. It may be large and attenuated.
A third is applying the attenuation formula to errors that are not classical. When the error is related to the true value, the bias can go either way.
Read the true variance.
$\sigma^2_{x^*} = 9$
Real variation.
Read the noise variance.
$\sigma^2_e = 1$
Reporting error.
Add the two variances.
$9 + 1 = 10$
The observed variance.
Divide true by observed.
$9 \div 10 = 0.9$
Reliability ratio.
Interpret the ratio.
$90\% \text{ signal}$
Mild attenuation to come.
Read the true slope.
$\beta_1 = 1.2$
What we want.
Read the variances.
$\sigma^2_{x^*} = 1, \ \sigma^2_e = 1$
Half the variance is noise.
Compute the reliability.
$1 \div 2 = 0.5$
Very noisy.
Shrink the slope.
$0.5 \times 1.2 = 0.6$
What OLS tends to.
Compute the bias.
$0.6 - 1.2 = -0.6$
Half the effect is lost.
State the direction.
$\text{toward zero}$
Same sign, smaller size.
Read the true slope.
$\beta_1 = -0.5$
A negative effect.
Read the variances.
$\sigma^2_{x^*} = 16, \ \sigma^2_e = 4$
From a validation study.
Add the two variances.
$16 + 4 = 20$
Observed variance.
Compute the reliability.
$16 \div 20 = 0.8$
Signal share.
Shrink the slope.
$0.8 \times (-0.5) = -0.4$
Still negative.
Compute the bias.
$-0.4 - (-0.5) = 0.1$
Positive: toward zero from below.
Check the direction.
$|{-0.4}| < |{-0.5}|$
Attenuation shrinks the size.
Write the correction.
$\beta_1 \approx \hat\beta_1 \div \lambda$
Undo the shrinkage.
Substitute and divide.
$0.6 \div 0.75 = 0.8$
Corrected slope.
Check the size.
Reported income has variance $100$ (in squared thousands of dollars) across households. A validation study against tax records finds that $70$ of that is true variation in income and the rest is reporting noise. What is the reliability ratio of reported income?
Answer: reliability ratio
Complete the worked solution: reported schooling has variance $10$, of which $6$ is true variation. A regression on reported schooling gives a slope of $0.216$. Find the noise variance, the reliability ratio and the corrected slope.
Subtract the true variance.
$\sigma^2_e = 10 - 6 =$ a
What the reports add on top of the truth.
Divide true by observed.
$\lambda = 6 \div 10 =$ b
The reliability ratio.
Divide the slope by the ratio.
$\hat\beta_1 \div \lambda =$ c
The corrected slope.
A regression using a noisily measured regressor gives a slope of $0.6$. The regressor's reliability ratio is known to be $0.5$. What is the corrected estimate of the true slope?
Answer: corrected slope
True regressor variance $4$, noise variance $1$, true slope $2$. Fill in the reliability ratio, the attenuated slope and the bias.
| value | |
|---|---|
| reliability ratio | |
| attenuated slope | |
| bias |
Match each measurement problem to its main consequence for the OLS slope.
| slope unbiased, standard error larger | slope biased toward zero | slope biased in a direction that depends on the pattern | |
|---|---|---|---|
| test scores recorded with random grading noise (the outcome) | |||
| years of schooling misreported at random (the regressor) | |||
| income understated more by high earners (the regressor) | |||
| hours worked recorded with random rounding (the outcome) |
The true regressor $x^*$ has variance $16$, but it is recorded with classical measurement error of variance $4$. The true slope is $-0.5$. What slope will OLS on the recorded $x$ tend to in large samples?
Answer: attenuated slope
Surveys of twins let labor economists compare one twin's report of schooling with the other twin's report of it. Such studies put the reliability of self-reported schooling near $0.9$. If a log-wage regression on reported schooling gives $0.72$, what is the slope corrected for measurement error?
Answer: corrected slope
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
True regressor variance $16$, noise variance $4$, true slope $-0.5$. Fill in the reliability ratio, the attenuated slope and the bias.
| value | |
|---|---|
| reliability ratio | |
| attenuated slope | |
| bias |
You can reason about noisy data. Explain to someone why random errors in a regressor do not average out in a large sample.
17. Your turn: observed slope 0.6, reliability 0.75., step 3
$0.8 > 0.6$
Correction always enlarges it.