Back to the on-screen lesson ·
Writing down what is on trial, choosing a one-sided or two-sided alternative from the question, and fixing the significance level before the data arrive.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will state a null and an alternative hypothesis for a research question, choose between a one-sided and a two-sided alternative from the question rather than from the data, split a significance level between the tails it has to cover, and name the decisions a test must settle before any data are collected.
A confidence interval asks what values are plausible. A test asks is this particular value plausible. The standard error, the critical value and the sampling distribution are all the same objects as in the last unit; what changes is that one value is singled out in advance and the data are asked whether they can embarrass it.
The null hypothesis is the statement of no effect, written about a population parameter and containing an equality. The alternative is what the study is looking for, and it is one-sided or two-sided according to whether a direction was specified. The significance level is the probability of rejecting a true null, chosen before the data. A Type I error is rejecting a true null.
A test names a specific value of a parameter and asks whether the data are compatible with it. The two statements are
$$H_0: \theta = \theta_0 \qquad \text{against} \qquad H_1: \theta > \theta_0, \;\; \theta < \theta_0 \;\; \text{or} \;\; \theta \ne \theta_0.$$
Three features of the null are worth stating out loud. It concerns the parameter, never the statistic — a hypothesis about the sample mean would be a hypothesis about a number everyone can see. It contains the equality, because the test must compute probabilities under one specific distribution and only an equality supplies one. And it is the statement being put on trial, so the burden of proof lies against it: the test can convict it, never acquit it.
The significance level $\alpha$ is fixed before the data. It is the probability, computed under $H_0$, that the test rejects — the rate at which an innocent null is convicted. A two-sided alternative splits $\alpha$ between the two tails; a one-sided alternative spends all of it in one, which makes the one-sided test more powerful in the direction it was pointed and blind in the other.
Everything before the data must happen before the data. Choosing the direction to match the sample, or moving the level to whichever side of the P-value is convenient, leaves the arithmetic correct and the error rate a fiction. That is the whole reason a test is described as a procedure rather than as a calculation.
Another way: story
A court presumes innocence, hears evidence and either convicts or fails to convict. It never declares the accused innocent. The null hypothesis is the presumption, the significance level is how strong the evidence must be, and fail to reject is exactly not proven — which is why a test that does not reject has established nothing at all about the null being true.
Another way: steps
A label says the mean weight is $500$ grams. What the study is looking for decides the alternative, and the alternative decides the rejection region.
| The question | Alternative | Rejection region at $5\%$ |
|---|---|---|
| Is the mean above $500$? | $\mu > 500$ | $z > 1.645$ |
| Is the mean below $500$? | $\mu < 500$ | $z < -1.645$ |
| Has the mean changed at all? | $\mu \ne 500$ | $|z| > 1.960$ |
The two-sided critical value is the largest of the three, because the same five per cent has to cover two tails instead of one. A one-sided test is therefore easier to pass — which is exactly why the direction must be fixed by the question rather than by the data.
Stating the alternative as the null. The null carries the equality and the presumption. Swapping them reverses which error the level controls.
Writing a hypothesis about the sample. $H_0: \bar x = 500$ is not a hypothesis; $\bar x$ is a number in the data.
Choosing the direction after seeing the sample. This doubles the true Type I error rate while the reported one stays the same.
Moving the significance level to suit the P-value. A level chosen afterwards controls nothing, and the number reported beside it is no longer the rate of anything.
A supplier guarantees a mean tensile strength of at least $70$; a buyer suspects otherwise.
The buyer is looking in one direction.
$H_0: \mu = 70$ against $H_1: \mu < 70$, at $\alpha = 0.05$.
Equality in the null.
Reject only for sufficiently small sample means; a large one, however extreme, is no evidence for this alternative.
One tail, chosen by the question.
A process is monitored for any change from its target of $12$.
No direction is of interest.
$H_0: \mu = 12$ against $H_1: \mu \ne 12$, at $\alpha = 0.05$.
Both tails matter.
Each tail gets $2.5\%$, so the critical value is $1.96$ rather than $1.645$.
The price of watching both directions.
The alternative is one-sided: the probability of heads exceeds a half.
Suspicion has a direction.
At the $5\%$ level the whole five per cent goes in the upper tail.
One tail only.
So an unusually low number of heads, however extreme, is not evidence for this alternative at all.
Each study concerns a population mean whose claimed value is $82$. Match each question to the alternative hypothesis it calls for.
| The mean is greater than $82$ | The mean is less than $82$ | The mean is not equal to $82$ | The difference between the two method means is not zero | |
|---|---|---|---|---|
| A regulator asks whether the mean dose is higher than the label says | ||||
| A consumer group asks whether the mean weight falls short of the label | ||||
| An engineer asks whether the process mean has shifted at all | ||||
| A laboratory asks whether two measuring methods read differently on average |
A claim that a population mean equals $80$ is to be tested. Put the steps of the test into the order they are carried out.
Number the steps in order (write the number in the box):
A test is run at the $8\%$ significance level. Give the percentage in each tail if the alternative is two-sided, and the percentage in the one tail if it is one-sided.
Two-sided, each tail: k. One-sided, the one tail: b.
A manufacturer claims that the mean life of its lamps is $540$ hours, and a consumer group suspects the true mean is lower. Which statement is the null hypothesis?
A researcher describes a study of a mean claimed to be $45$ in four sentences. Mark the two decisions that had to be made before the data and were not.
This task has no paper form; do it on a device.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A test is built to reject a true null hypothesis $8\%$ of the time. What is its Type I error rate, as a proportion?
Answer:
You can write a null and an alternative for a claim, say which contains the equality and why, and identify a decision that was made too late. Say in your own words why the burden of proof lies against the null.
9. Your turn: a coin is suspected of favouring heads, and $H_0$ is that it is fair, step 3