Back to the on-screen lesson ·

Designing a study

Populations and samples, the sampling methods and their biases, observational studies against experiments, confounding, control, blinding — and what each design allows a conclusion to say.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you can look at how a set of data was produced and say what it is allowed to be evidence for. You can name the sampling methods and the biases each is exposed to, tell an observational study from an experiment, explain what a confounding variable is and why more data never removes one, and state the strongest claim a given design supports — including, when it matters, that the answer is none.

2. What you bring to this

You can summarise a set of numbers. This lesson asks the question that comes before that one and decides whether the summary means anything: where did the numbers come from, and what are they allowed to be evidence for?

3. Words you will need

Population: everyone or everything the conclusion is about.

Sample: the part actually measured.

Bias: a design that misses in the same direction every time. More data does not reduce it.

Observational study: the researcher watches. Experiment: the researcher assigns.

Confounding variable: something that moves with both the supposed cause and the effect, so it could explain the pattern by itself.

Blinding: keeping participants, or assessors, from knowing which group somebody is in.

4. The case: does a night light cause short-sightedness?

In 1999 a study in Nature reported that children who had slept with a night light as babies were far more likely to be short-sighted. The association was strong, the sample was large, and the finding travelled around the world as advice to parents.

The case for the conclusion. Over 470 children were surveyed. The pattern was clear and graded: the brighter the light, the higher the rate. A plausible mechanism existed — light during sleep might affect how the eye grows.

The case against it. Nobody assigned any child to a night light. Parents chose, and short-sighted parents — whose children inherit the tendency — are more likely to install one, because they cannot see well in the dark themselves. A follow-up study measured the parents' eyesight, and the association between night lights and children's eyesight disappeared.

What decided it. Not the size of the sample, and not the strength of the association. It was the design. An observational study can establish that two things go together; only random assignment can rule out the explanations nobody thought of — including the one that turned out to be right here.

5. What a design lets you claim

Every conclusion drawn from data rests on how the data was produced, and two questions settle what may be claimed.

Were the units randomly selected from the population? If they were, the sample has no systematic reason to differ from the population, and what is found in it can be generalised. If they selected themselves — a phone-in, a comment card, an online poll — the sample describes the people who volunteered and nobody else, and no amount of data repairs that.

Were the units randomly assigned to the treatments? If they were, the groups are alike in every variable, including the ones nobody measured or imagined, so a difference in outcome is attributable to the treatment. If not — if people chose, or were sorted by anyone's judgement — then any variable behind that choice is an alternative explanation, and the study cannot rule it out.

Sampling methods differ in how the selection is organised: simple random, stratified (groups first, then sample within each), cluster (whole groups chosen at random), systematic (every $k$th). Convenience sampling is not one of these; it is the absence of a method. And within an experiment, control, randomisation, replication and blinding are the four features that make the comparison mean something.

Another way: picture

Picture two doors into a study. The first door is who gets in — random selection widens the conclusion from the room to the world. The second door is which group they go to — random assignment turns a difference between the groups into a claim about the treatment. Walk through one door and you have half a conclusion.

Another way: steps

  1. Name the population and the sample, and ask who could not have been chosen.
  2. Ask whether anybody was assigned, or whether the researcher watched.
  3. If assigned: was it at random, and was anyone blinded?
  4. Write the strongest claim the design supports — and say which of generalisation and causation it does not support.

6. What each design supports

DesignMay generaliseMay claim a cause
Random sample, observedYesNo
Volunteers, observedNoNo
Random sample, randomly assignedYesYes
Volunteers, randomly assignedOnly to people like themYes

The last row is most experiments: a causal conclusion about a group nobody chose at random. That is a real result and a limited one, and saying so is part of reporting it.

7. Three things that trip people up

"A big enough sample fixes a bad sample." It does the opposite: it makes a biased estimate more precise, so the wrong answer arrives with a narrower interval around it. A poll of a million self-selected readers is worse evidence than a random sample of a thousand.

"Randomised means random selection." There are two randomisations and they do different jobs. Random selection from the population lets you generalise. Random assignment to treatments lets you claim a cause. A study can have either, both, or neither.

"Statistically significant means important, or means causal." It means the result would be unlikely by chance alone. In an observational study a significant result is a significant association, and every confounding variable is still standing exactly where it was.

8. A drug trial with a placebo group

  1. Volunteers are assigned at random to the drug or an identical dummy pill.

    Random assignment: the groups start alike.

  2. Neither the patients nor the doctors assessing them know which is which.

    Double blinding: expectation cannot favour one group.

  3. A difference in recovery is attributable to the drug — for people like these volunteers. Generalising further needs an argument the data cannot supply.

    Causation yes, generalisation with care.

9. A survey of a school's students

  1. Ten students are drawn at random from each year group.

    Stratified sampling: every year is represented.

  2. They are asked how many hours they study. Nobody is assigned anything.

    Observational.

  3. The average can be reported for the school. If studious students are also the ones who reply, even that is in doubt — non-response is a bias that survives random selection.

    Selection is necessary, not sufficient.

10. Guided practice

To estimate the average commute for the workers of a city, which sample should be used?

11. Guided practice

A factory tests every fiftieth item coming off the line. What sampling method is this?

12. Practice

An observational study finds that people who drink more coffee sleep less. Why can we not conclude that coffee causes poor sleep?

13. Practice

Match each description to the method it describes.

StratifiedClusterSystematicConvenience
Sample within every group
Choose whole groups, survey all of them
Every fiftieth item on the line
Whoever walks past

14. Practice

Put the stages of a well-designed experiment in order.

Number the steps in order (write the number in the box):

15. Practice

Select the parts of this survey question that push the reader toward an answer.

This task has no paper form; do it on a device.

16. Somewhere new

A newspaper reports: “In a study of 40,000 people, towns with more fire engines have larger fires. Researchers say the effect was large and highly significant.” What is the strongest fair headline?

17. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

18. Test question

A trial draws its participants at random from the households of a county but lets each of them choose which of the two programmes to follow. What may its result be used for?

19. What you can do now

You can judge a study design and say what it licenses. Without looking: which randomisation lets you generalise, which lets you claim a cause, and why does a larger sample not fix a biased one?