Back to the on-screen lesson ·
Comparing observed counts with what a model expects, goodness of fit and tests of independence, and inference for the slope of a least-squares line.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you can test whether counts in categories fit a claimed distribution, and whether two categorical variables are associated, computing expected counts, contributions and degrees of freedom. You can also treat a regression slope as the estimate it is: compute its $t$ statistic on $n - 2$ degrees of freedom, build an interval for the true slope, and check the residual plot that decides whether any of it is trustworthy.
You can run a test on a mean or a proportion — one quantity, measured. This lesson handles data that is not measured at all but counted into categories, where the question is whether the counts fit a claim or whether two classifications are related.
Observed count $O$: what the data holds.
Expected count $E$: what the null hypothesis predicts.
Goodness of fit: do counts across one variable's categories match a claimed distribution?
Test of independence: are two categorical variables associated?
Degrees of freedom: how many cells are free once the totals are fixed.
Both chi-square tests compare observed counts with the counts a null hypothesis would predict, using the same statistic:
$$\chi^2 = \sum \frac{(O - E)^2}{E}$$
Squaring makes every disagreement positive; dividing by $E$ makes them comparable, so that being five out matters more in a cell expecting ten than in one expecting a thousand. Cells that match contribute nothing, so the statistic is a running total of disagreement and is never negative.
Goodness of fit asks whether counts across one variable's categories match a claimed distribution. $E$ is the claimed proportion times $n$, and $df = k - 1$.
Independence asks whether two classifications are related. Under the null they are not, so $E = \dfrac{\text{row total} \times \text{column total}}{n}$, and $df = (\text{rows} - 1)(\text{columns} - 1)$.
The conditions matter: the data must be counts, not percentages or means; the observations must be independent; and every expected count should be at least about $5$, because the chi-square distribution is an approximation that fails in very small cells.
Another way: picture
Picture two tables side by side: the counts that were seen, and the counts independence would have produced. The statistic walks the cells in step, squares each gap, and scales it by how large the cell should have been. A table that matches gives zero; every disagreement pushes the total up and never down.
Another way: steps
"Chi-square works on percentages." It works on counts. Feeding it percentages throws away the sample size, and the sample size is the whole reason a difference is or is not convincing.
"Degrees of freedom are the number of cells." They are the number of cells free to vary once the totals are fixed: $k - 1$ for one row, $(\text{rows}-1)(\text{columns}-1)$ for a table.
"A significant chi-square shows the variables are causally linked." It shows the counts are hard to reconcile with independence. Nobody was assigned to a category, so everything in the study-design lesson applies unchanged.
$60$ rolls, so each face expects $10$. Observed: $8, 12, 9, 11, 14, 6$.
Equal expected counts.
Contributions: $0.4, 0.4, 0.1, 0.1, 1.6, 1.6$, totalling $\chi^2 = 4.2$.
$(O-E)^2/E$ per face.
$df = 5$, and the $5\%$ critical value is $11.07$. $4.2$ is well below it: no evidence the die is unfair.
Not evidence that it is fair, either.
Row totals $60$ and $40$, column totals $50$ and $50$, $n = 100$.
The margins.
Expected: $60 \times 50 / 100 = 30$, and likewise $30, 20, 20$.
Row times column over $n$.
$df = (2-1)(2-1) = 1$, so the $5\%$ critical value is $3.841$.
One degree of freedom.
Each category expects $80/4 = 20$.
Contributions $\frac{4}{20} + \frac{4}{20} + 0 + 0 = 0.4$.
Matching cells contribute nothing.
$df = 3$, critical value $7.815$: nowhere near. The counts are exactly what equal likelihood predicts, give or take.
In a two-way table of $150$ people, a row totals $90$ and a column totals $30$. If the two variables were unrelated, what count would that cell be expected to hold?
answer
A cell holds $48$ when $45$ was expected. What does it contribute to the $\chi^2$ statistic?
answer
A goodness-of-fit test compares observed counts across $5$ categories with what a model predicts. How many degrees of freedom does it have?
answer
$60$ observations are spread over $6$ categories, and a model says all $6$ are equally likely. The observed counts are $15, 5, 10, 10, 10, 10$. Find $\chi^2$, to three decimal places.
answer
A report cross-classifies $100$ people two ways, finds $\chi^2 = 0.889$ on one degree of freedom with $p = 0.3458$, and concludes that the first variable causes the second. What is wrong?
You can fit a least-squares line and read its slope in context, and you can build a confidence interval and run a test. This lesson puts the two together: the slope from one sample is an estimate, and it has a standard error like any other estimate.
$\beta$: the true slope in the population — unknown.
$b$: the slope computed from this sample — an estimate of $\beta$.
$SE_b$: how much $b$ would vary from sample to sample.
$H_0: \beta = 0$: the hypothesis that $x$ carries no linear information about $y$.
$n - 2$ degrees of freedom: two coefficients were estimated, not one.
A different sample would have given a different line. So the computed slope $b$ is an estimate of an unknown population slope $\beta$, with a standard error $SE_b$ — and everything from the intervals and tests lessons applies to it unchanged.
The test. $H_0: \beta = 0$ says $x$ carries no linear information about $y$. The statistic is the usual count of standard errors, $t = \dfrac{b}{SE_b}$, on $n - 2$ degrees of freedom — two, because the line estimated both a slope and an intercept.
The interval. $b \pm t^* SE_b$, on the same degrees of freedom. It is the more useful report, because it says how large the slope plausibly is rather than only that it is not zero. If it contains zero, the test would not have rejected.
The conditions, which come first. The relationship must be linear, the residuals must have roughly constant spread across $x$ and be roughly normal, and the observations must be independent. All of these are read off the residual plot, and a curve or a funnel there invalidates the inference no matter how small the $p$-value is. Checking the plot after a significant result — rather than before — is how a curved relationship gets reported as a linear one.
Another way: picture
Picture ten different samples from the same population, each fitted with its own line, all drawn on one axis. They fan out around the true relationship. $SE_b$ measures how wide that fan is, and the test asks whether a horizontal line would fit inside it.
Another way: steps
| Estimate | Standard error | Degrees of freedom |
|---|---|---|
| $\bar{x}$ | $s/\sqrt{n}$ | $n - 1$ |
| $\hat{p}$ | $\sqrt{\hat{p}(1-\hat{p})/n}$ | — (normal) |
| $b$ | $SE_b$ | $n - 2$ |
Estimate $\pm$ critical value $\times$ standard error, every time. Only the middle column and the degrees of freedom change, and both are decided by how many things the data had to estimate.
"A significant slope means the line is a good model." It means the slope is unlikely to be zero. A curve fitted with a straight line can easily have a significant slope and a residual plot that is obviously wrong.
"Degrees of freedom are $n - 1$." They are $n - 2$ here, because the line has two estimated coefficients. Using $n - 1$ makes every interval slightly too narrow.
"A significant slope shows $x$ causes $y$." It shows an association unlikely to be chance. Regression is arithmetic on observed data; it cannot manufacture the random assignment that a causal claim needs.
$n = 20$ points, $b = 8$, $SE_b = 2$.
From the output.
$t = 8 / 2 = 4$ on $18$ degrees of freedom.
Count standard errors.
That is far into the tail, so reject $H_0$: there is evidence of a real linear association.
Association, not cause.
$b = 4$, $SE_b = 2$, $t^* = 2.228$ on $10$ degrees of freedom.
Margin $= 2.228 \times 2 = 4.456$, so the interval runs from $-0.456$ to $8.456$.
It contains zero, so a flat relationship is still plausible and the test would not have rejected — the same conclusion, with the size of the effect shown as well.
$df = 27 - 2 = 25$.
Two coefficients estimated.
$t = 15 / 5 = 3$, which is well into the tail.
Reject $H_0$ — but only after the residual plot has been checked, because none of this means anything if the shape was wrong.
A regression on $12$ points gives slope $b = 10$ with standard error $SE_b = 5$. Find the $t$ statistic for $H_0: \beta = 0$.
answer
A least-squares line is fitted to $20$ points. How many degrees of freedom does inference about its slope have?
answer
A regression on $20$ points gives $b = 4$ with $SE_b = 2$, and $t^* = 2.101$ for $95\%$ confidence on $18$ degrees of freedom. Give the lower and then the upper end of the interval for the true slope.
from lo to hi
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A test gives $\chi^2 = 2.5$ on $1$ degrees of freedom, with $p = 0.1138$. Which question was being asked?
A regression on $20$ points gives a slope significantly different from zero, but its residual plot shows a clear curve. What follows?
You can run chi-square tests and inference on a slope. Without looking: why does each cell's contribution divide by the expected count, and why does a regression lose two degrees of freedom rather than one?
8. Your turn: $80$ observations over $4$ equally likely categories, observed $22, 18, 20, 20$, step 3
21. Your turn: $n = 27$, $b = 15$, $SE_b = 5$, step 3