Back to the on-screen lesson ·

Test statistics and rejection regions

Turning a gap into a count of standard errors, drawing the rejection region before the data, and seeing that the null values a test spares are a confidence interval.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will compute a test statistic as a gap divided by a standard error, translate a rejection region between the statistic's scale and the original units, and find the set of null values a test would not reject. You will also say what a test statistic does not measure.

2. From a gap to a number that can be compared

The last lesson wrote down what is being tested. This one turns a difference measured in grams, hours or pounds into a number with no units at all, so that one table of critical values can serve every question anybody asks. The ingredient is the standard error from the first unit.

3. Words for this lesson

A test statistic measures the departure of an estimate from the null value in units of its own standard error. The null distribution is the distribution of that statistic when the null hypothesis is true. The rejection region is the set of values of the statistic for which the test rejects, and the critical value is where it begins.

4. One number, no units

For a mean with a known standard error, the statistic is

$$z = \frac{\bar x - \mu_0}{\text{standard error}},$$

the gap between what was estimated and what was hypothesised, divided by how far such an estimate typically lands from the truth. The division does two things. It removes the units, so that a difference in grams and a difference in kilometres become comparable. And it supplies a null distribution — standard normal here — from which probabilities can be read.

The rejection region is chosen before the data, from that distribution and the significance level. For a two-sided test at level $\alpha$ it is $|z| > z_{\alpha/2}$; translated back into the original units, it is every sample mean outside

$$\mu_0 \pm z_{\alpha/2} \times \text{standard error}.$$

Turn that inequality around and something familiar appears. The null values a test would not reject are exactly those within $z_{\alpha/2}$ standard errors of the estimate — which is the confidence interval of the last unit. A test and an interval are the same statement read in opposite directions: the interval answers every test at once, and a test asks whether one value is in the interval.

What the statistic is not is worth as much as what it is. It is not a probability; it can be any real number. It is not the size of the effect; a negligible effect measured very precisely gives a huge statistic. It says how surprising the data are, and says nothing at all about whether anyone should care.

Another way: picture

The null distribution drawn as a bell centred on zero, with the rejection region shaded in both tails beyond the critical values. The observed statistic is a single vertical mark. The test is nothing more than asking whether the mark landed in the shading — and the shading was painted before the mark was made.

Another way: steps

  1. Compute the gap: estimate minus null value.
  2. Divide by the standard error to get the statistic.
  3. Read the critical value for the level and the alternative.
  4. Compare, and report the decision with the statistic beside it.

5. The same gap, four standard errors

A gap of $6$ between the estimate and the null value, at four different precisions.

Standard errorStatisticReject at $5\%$, two-sided?
61.0no
32.0yes
23.0yes
16.0yes

The gap never changed. Everything in the last column was decided by how precisely the gap was measured, which is to say by the sample size. That is why a large study can find a trivial difference significant and a small one can miss a large difference entirely — and why the gap itself has to be reported beside the statistic.

6. Where this goes wrong

Dividing by the standard deviation of the data instead of the standard error of the estimate. This understates the statistic by a factor of $\sqrt n$, and the error grows with the sample size.

Reading the statistic as a probability. It is a distance. A value of $3$ is not a probability of three.

Reading a large statistic as a large effect. Precision alone can produce one. Report the estimate and its interval beside the statistic and the confusion cannot survive.

Comparing a two-sided statistic with a one-sided critical value. The critical value depends on the alternative, and mixing them silently doubles or halves the error rate.

7. A statistic and a decision

  1. $H_0: \mu = 100$; $\bar x = 106$, standard error $2.5$, two-sided at $5\%$.

    Everything fixed in advance.

  2. $z = (106 - 100)/2.5 = 2.4$.

    Gap over standard error.

  3. $|2.4| > 1.96$, so the null is rejected — and the estimated difference of $6$ is reported beside it.

    The decision and the size, together.

8. The rejection region in the original units

  1. The same test: reject when $\bar x$ is outside $100 \pm 1.96 \times 2.5$.

    Translate the critical value back.

  2. That is outside $(95.1, 104.9)$.

    Two boundaries, fixed before the data.

  3. The observed $106$ is outside, which is the same decision reached the same way.

    Two routes, one answer.

9. Your turn: $H_0: \mu = 50$, $\bar x = 44$, standard error $4$, two-sided at $5\%$

  1. The gap is $44 - 50 = -6$.

    Estimate minus null value.

  2. The statistic is $-6/4 = -1.5$.

    Divide by the standard error.

  3. Your turn: work this step out. Its working is at the end of the packet.

    $|-1.5| < 1.96$, so the test does not reject.

10. Guided practice

The null value of a population mean is $57$, the sample mean is $63$, and the standard error of the sample mean is $2$. What is the test statistic?

Answer:

11. Guided practice

A two-sided test of the claim that a population mean is $55$ uses a sample mean whose standard error is $1$. Give the two boundaries of the rejection region, in the original units, at each significance level.

Lower boundaryUpper boundary
At the $10\%$ level
At the $5\%$ level
At the $1\%$ level

12. Practice

A sample mean of $57$ is compared with a null value of $52$, and its standard error is $2$. Give the gap in the original units, and the test statistic.

Gap: g. Statistic: z.

13. Practice

A sample mean of $32$ has standard error $4$. Give the set of null values that a two-sided test at the $20\%$ significance level would fail to reject, with both endpoints included.

This task has no paper form; do it on a device.

14. Somewhere new

Four studies test the same two-sided null hypothesis and report the statistics below. Put them in order of the evidence they carry against the null, weakest first.

Number the steps in order (write the number in the box):

15. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

16. Test question

A test statistic comes out at $6$. What does that number measure?

17. What you can do now

You can compute a test statistic, mark out a rejection region before any data, and recognise the confidence interval hiding inside a two-sided test. Say in your own words why a large statistic is not the same thing as a large effect.

Working for the steps left to you

9. Your turn: $H_0: \mu = 50$, $\bar x = 44$, standard error $4$, two-sided at $5\%$, step 3