Back to the on-screen lesson ·
Test several coefficients at once from sums of squares or R-squared, run the overall F test, and see why a group can matter when no member does alone.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to count restrictions, compute an F statistic from SSR or R-squared, and interpret a joint test.
You can test one coefficient with a t statistic and you know the sum of squared residuals measures how badly a regression fits. From the last two lessons you know why controls are added and that adding any regressor lowers SSR at least a little.
Many questions are about a group of coefficients. Do family-background variables matter for wages once schooling is controlled? Do three region dummies add anything? Is the whole regression better than nothing? Testing each coefficient separately answers the wrong question, and this lesson shows the test that answers the right one.
| Term | What it means |
|---|---|
| Joint hypothesis | A null that restricts several coefficients at once, such as $\beta_3 = \beta_4 = 0$. |
| Restriction | One equation in the null; $q$ is their number. |
| Unrestricted model | The regression that includes every variable in question. |
| Restricted model | The regression with the null imposed, for example the variables dropped. |
| F statistic | The loss of fit per restriction divided by the unrestricted error variance. |
| Overall F test | The joint test that every slope in the regression is zero. |
| Exclusion restriction | A null that a group of variables can be dropped from the model. |
Fit the regression twice: once freely (unrestricted) and once with the null imposed (restricted — usually by dropping the variables). Imposing restrictions can only make the fit worse, so $SSR_r \ge SSR_{ur}$. The question is whether it gets worse by more than chance would explain.
The F statistic scales the loss of fit by the noise level:
$$F = \frac{(SSR_r - SSR_{ur})/q}{SSR_{ur}/(n - k - 1)}.$$
The numerator is the average loss per restriction; the denominator is the unrestricted estimate of the error variance. Under the null, and the usual assumptions, $F$ has an F distribution with $q$ and $n - k - 1$ degrees of freedom. Large values — beyond the 5 percent critical value, typically around $2$ to $4$ for the sample sizes economists use — reject the null.
Since $SSR = (1 - R^2)SST$ and both models share the same SST, the statistic can also be written with $R^2$:
$$F = \frac{(R^2_{ur} - R^2_r)/q}{(1 - R^2_{ur})/(n - k - 1)}.$$
Another way: action
Drop a group of variables from a regression and watch SSR rise. If it rises a lot compared with the noise in the model, those variables were doing real work together.
Another way: steps
Suppose you test three coefficients separately, each at the 5 percent level. Even if all three are truly zero, the chance that at least one looks significant is about $14$ percent, not $5$. Separate tests do not control the chance of a false alarm for the group.
The opposite problem is just as common. When regressors are highly correlated — mother's and father's education, say — each coefficient can have a large standard error, because the data cannot tell which parent's schooling matters. Neither t statistic is significant. Yet dropping both variables may hurt the fit a great deal, because together they clearly matter. The F test measures exactly that joint contribution and is not confused by how it is shared between the variables.
So a group can be jointly significant when no member is, and a single member can look significant by chance when the group is not. The question asked should decide the test.
The rise in SSR from imposing true restrictions is pure noise: each restriction costs, on average, about one error variance's worth of fit, because the unrestricted model was able to chase a little noise with each extra coefficient. So under the null, the numerator is an estimate of $\sigma^2$. The denominator is another estimate of $\sigma^2$. Their ratio should be near one.
If the restrictions are false, the variables genuinely explain $y$, and dropping them costs far more than noise. The numerator grows while the denominator does not, and $F$ becomes large. That is why the test rejects only for large $F$: small values are exactly what the null predicts. With a single restriction, $F$ equals the square of the t statistic, so the two tests agree.
The regression output of every statistics package reports an F statistic for the null that all slopes are zero. The restricted model has only an intercept, so its $R^2$ is zero and the formula reduces to
$$F = \frac{R^2/k}{(1 - R^2)/(n - k - 1)}.$$
Rejecting says the regressors together explain more than nothing. That is a low bar, and in large samples almost every regression clears it, even with $R^2$ of a few percent. It is useful mainly as a sanity check. The tests that answer economic questions are about specific groups of variables chosen for a reason.
Not every joint hypothesis drops variables. Economic theory often predicts that coefficients take particular values or relate to each other. A Cobb-Douglas production function with constant returns to scale predicts that the coefficients on log capital and log labor add up to one. A tax study might ask whether a dollar of wage income and a dollar of capital income have the same effect on spending.
The F test handles these the same way. Impose the restriction by rewriting the model — for constant returns, regress log output per worker on log capital per worker — fit the restricted model, and compare its SSR with the unrestricted SSR. Each equation in the null is one restriction. The only care needed is that the restricted model has the same dependent variable, or the SSR form must be used rather than the R-squared form.
This is how theory meets data in much of applied economics. A theory makes restrictions; an unrestricted regression lets the data speak freely; the F test asks whether imposing the theory costs more fit than chance would explain. When it does, the theory is rejected in that form, and the size of the loss suggests how far off it is.
The F distribution has two degrees-of-freedom numbers: $q$ for the numerator and $n - k - 1$ for the denominator. Its 5 percent critical value falls as either grows. With one restriction and a large sample it is about $3.84$, the square of $1.96$. With two restrictions it is about $3.0$; with three, $2.6$; with five, $2.2$; and with ten, $1.8$.
Those values have a simple reading. Under the null, $F$ averages about one, since both numerator and denominator estimate the error variance. The more restrictions are tested together, the more the numerator averages over, the less it wanders from one by chance, and the smaller the value needed to be surprising.
In practice, software reports the p-value directly: the probability, if the null were true, of an F at least as large as the one observed. As with t tests, a p-value below $0.05$ means rejection at the 5 percent level. And as with t tests, rejection says the restrictions are false, not that the variables matter a great deal. With a huge sample, a trivially small joint contribution can still produce a large F, so the size of the loss of fit — how much $R^2$ falls — belongs in the report alongside the test, so a reader can judge whether the variables matter in practice as well as in principle, which is what a decision needs.
A regression of log wages on schooling, experience and two parental-education variables has $R^2 = 0.30$ with $70$ residual degrees of freedom. Dropping the parental variables lowers $R^2$ to $0.20$.
Three checks.
Then state the conclusion about the group, not about any one coefficient.
Labor economists often ask whether family background affects earnings once a person's own schooling and experience are accounted for. A typical approach adds the mother's and father's years of schooling to a log-wage regression and tests whether both coefficients are zero.
Because the two parents' schooling levels are strongly correlated — people tend to marry people with similar education — each coefficient is imprecise, and the regression cannot separate the two parents' contributions with any confidence. In one sample, the t statistics might be $1.4$ and $1.2$, neither significant. But dropping both lowers $R^2$ from $0.36$ to $0.28$ with $64$ residual degrees of freedom, giving $F = 4$, above the critical value of about $3.1$. Family background matters jointly, even though the data cannot say which parent's schooling carries the effect.
The same test is used in wage-discrimination cases, marketing studies and program evaluations whenever a question concerns a set of variables: region dummies, industry dummies, a set of policy instruments. Asking the question about the group, with one test, is both statistically correct and closer to what the decision maker wants to know. It also protects against the temptation to keep only the one dummy that happened to come out significant, a habit that turns noise into apparent findings that later studies then fail to reproduce.
The most common mistake is dropping a group of variables because none is individually significant. Correlated variables can each be insignificant and jointly decisive.
A second mistake is putting the restricted model's degrees of freedom in the denominator. The denominator is the unrestricted model's error variance.
A third is reading a significant overall F as evidence that the regression is causal or even useful. It only says the regressors explain more than nothing.
Read the two sums.
$SSR_r = 240, \ SSR_{ur} = 200$
Dropping variables worsened fit.
Read the restrictions and df.
$q = 2, \ n - k - 1 = 100$
From the setup.
Compute the numerator.
$(240 - 200)/2 = 20$
Loss per restriction.
Compute the denominator.
$200/100 = 2$
Error variance.
Divide the two.
$F = 20/2 = 10$
Far above typical critical values.
Read the fit.
$R^2 = 0.2$
Twenty percent explained.
Read the model size.
$k = 4, \ n = 105$
Four regressors.
Count the degrees of freedom.
$105 - 4 - 1 = 100$
For the denominator.
Compute the numerator.
$0.2/4 = 0.05$
Fit per regressor.
Compute the denominator.
$0.8/100 = 0.008$
Noise per degree of freedom.
Divide the two.
$F = 6.25$
The regressors matter jointly.
State the null.
$H_0: \beta_{mom} = \beta_{dad} = 0$
Two restrictions.
Read the t statistics.
$t_{mom} = 1.4, \ t_{dad} = 1.2$
Neither above 1.96.
Read the R-squared values.
$R^2_{ur} = 0.36, \ R^2_r = 0.28$
Dropping both costs fit.
Compute the numerator.
$(0.36 - 0.28)/2 = 0.04$
Per restriction.
Compute the denominator.
$(1 - 0.36)/64 = 0.01$
With 64 degrees of freedom.
Divide the two.
$F = 4$
Above about 3.1.
Reconcile the results.
$\text{correlated regressors share the effect}$
Together they clearly matter.
Compute the numerator.
$(150 - 120)/3 = 10$
Loss per restriction.
Compute the denominator.
$120/60 = 2$
Error variance.
Divide the two.
An unrestricted regression has $SSR_{ur} = 300$ with $100$ degrees of freedom. Imposing $q = 5$ restrictions raises the sum of squared residuals to $SSR_r = 330$. What is the F statistic?
Answer: F statistic
Complete the worked solution: an unrestricted regression has $SSR_{ur} = 300$ with $100$ degrees of freedom. Imposing $q = 4$ restrictions raises SSR to $384$. Find the numerator, the denominator and F.
Divide the rise in SSR by q.
$84 \div 4 =$ a
Loss of fit per restriction.
Divide the unrestricted SSR by its df.
$300 \div 100 =$ b
Unrestricted error variance.
Divide the numerator by the denominator.
$F =$ c
Compare with the critical value.
A regression with $k = 4$ regressors on $n = 105$ observations has $R^2 = 0.2$. What is the F statistic for the null that every slope is zero?
Answer: overall F statistic
Restricted $SSR = 240$, unrestricted $SSR = 200$, $q = 2$ restrictions and $100$ unrestricted degrees of freedom. Fill in the numerator, the denominator and F.
| value | |
|---|---|
| numerator | |
| denominator | |
| F statistic |
Match each null hypothesis to its number of restrictions $q$.
| q = 1 | q = 2 | q = 3 | q = 4 | |
|---|---|---|---|---|
| β2 = 0 and β3 = 0 and β4 = 0 | ||||
| β2 = β3 | ||||
| all four slopes in a four-regressor model are zero | ||||
| β2 = 0 and β3 = 1 |
A log-wage regression with family-background variables has $R^2 = 0.25$ and $150$ residual degrees of freedom. Dropping the $1$ background variables lowers $R^2$ to $0.2$. What is the F statistic for the null that all $1$ background coefficients are zero?
Answer: F statistic
A retail chain regresses store sales on price, advertising and $5$ region dummies. The full regression has $SSR = 300$ with $100$ degrees of freedom; without the region dummies, $SSR = 330$. What is the F statistic for the null that regions do not matter?
Answer: F statistic
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Restricted $SSR = 330$, unrestricted $SSR = 300$, $q = 5$ restrictions and $100$ unrestricted degrees of freedom. Fill in the numerator, the denominator and F.
| value | |
|---|---|
| numerator | |
| denominator | |
| F statistic |
You can test several restrictions at once. Explain to someone how two variables can each be insignificant and still matter together.
17. Your turn: SSR_r = 150, SSR_ur = 120, q = 3, df = 60., step 3
$F = 10/2 = 5$
Compare with the critical value.