Back to the on-screen lesson ·

How the sample was chosen

The rule that decided who got into the sample, and why a bigger sample never repairs a rule that leaves a group out.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will read past a survey's number to the sentence that says how its people were picked, state that rule in one line, and name a group the rule shuts out. You will identify specific coverage and response risks, and explain why increasing the size of a subgroup sample does not automatically extend its coverage or remove bias.

2. Two surveys about the same bus route

Two surveys about the same bus route

The city council has to decide whether to keep the number 7 running in the evening, and two surveys have been handed in.

The first survey

Nine hundred passengers were asked, on the bus, whether the evening service should be kept.

Eight hundred and sixty said yes, which the report gives as 96 percent.

The second survey

Two hundred names were drawn from the electoral register for the three wards the route runs through.

A hundred and sixty of them answered, and ninety said the evening service should be kept.

What the covering note says

The note recommends the first survey, because it asked more than four times as many people.

3. What you already have

From lesson 1 you can name the two groups: the sample that was measured and the population the claim is about. This lesson is about the bridge between them — the rule that decided which members of the population ended up in the sample.

4. Words for this lesson

TermWhat it means
Probability sampleA design using a chance mechanism with known selection probabilities.
Convenience sampleCases selected because they are easy to reach.
Self-selectionParticipants deciding for themselves whether to enter the observed group.
NonresponseMissing replies from cases invited or selected for measurement.
Inference to the best explanationComparing candidate accounts of evidence and favoring the best-supported available account provisionally.

5. Which survey should the city council act on?

Read both surveys. The city council will keep the evening bus or take it away, and neither choice is free: somebody loses a way home, or the money goes on a route nobody rides.

The covering note recommends the larger survey, and that is not a silly recommendation. Nine hundred is more than a hundred and sixty.

But ask who each rule could reach: one was handed out on the bus, the other drawn from a register of everybody in the wards, ridden or not.

Decide before you read on.

6. Inspect the selection mechanism

A sample's value depends on the question and how cases entered it. A particular subgroup can be useful for studying that subgroup while providing a weak basis for a population-wide estimate. A large sample does not automatically repair a route that systematically misses relevant cases.

A fixed-size simple random sample gives every subset of that size an equal selection probability; drawing distinct names uniformly without replacement from a complete register is an introductory example. Probability sampling more broadly uses a documented chance mechanism with known nonzero inclusion probabilities. Nearby people, volunteers, and whoever answers first are not made random merely by the researcher's lack of a plan.

Ask how a case could enter, who was excluded, and whether selection or response plausibly relates to the outcome. This identifies a risk of bias rather than proving the sample's actual error or its direction. A probability design also needs suitable coverage, follow-up, and measurement. Good selection is one part of an inquiry, not a guarantee that every later step is flawless.

Another way: picture

Draw the target population, the available list, the selected cases, and the respondents as successive boundaries. Label each narrowing with its rule. The diagram distinguishes an incomplete list from selective response after a random invitation.

Another way: steps

State the target population. Identify the selection rule and any chance mechanism. Check coverage and response. Explain a specific connection that could make included and excluded cases differ on the outcome.

7. Random, systematic, convenient, and self-selected

A uniform random draw from a complete register gives each eligible person a specified chance of selection. The crucial feature is the chance mechanism and its known relationship to the list, not whether the result happens to look balanced. A small random sample can contain an unusually large share of one subgroup by chance. That possibility is sampling variation. It differs from a rule that deliberately or systematically selects only that subgroup.

A systematic sample takes every kth item after a starting point. A random start can give a useful probability-based design, but the ordering deserves inspection. If every tenth manufactured item comes from a different machine and you always inspect every tenth item in the same phase, the sample may track one machine rather than the whole process. A random start improves the selection mechanism, yet a strong repeating pattern can still make one realized sample unrepresentative. Do not define every-tenth selection as automatically random or automatically safe.

A convenience sample uses cases easy to reach. Asking people leaving a gym may be efficient for learning about those gym users. It is weaker for estimating the exercise habits of every resident in the district because nonusers are absent. The problem is the link between access and the outcome being measured. If the question concerns the gym's locker room layout, those same respondents might be exactly the relevant population. A method is assessed relative to its target question.

Self-selection occurs when people decide whether to enter the study, as with an open phone-in or optional website review. Strong feelings, available time, internet access, and interest in the topic can all affect participation. Those factors may also affect the answers. This creates a reason to investigate representativeness. It does not tell us automatically that the result is too high, too low, or dominated by both extremes. The direction and size of the difference need evidence rather than a stereotype about reviewers.

Nonresponse can occur after a good initial draw. If 100 randomly chosen people are invited but only 30 answer, the realized sample is not simply the original random list. Compare what is known about responders and nonresponders and consider suitable follow-up. High response is helpful, but a response percentage alone does not establish absence of bias. A small missing group might matter greatly if its answers differ systematically on the question of interest.

Report each stage rather than announcing a single magic label. 'We drew 100 names uniformly from the complete register, obtained 80 responses, and followed up missing replies by another route' communicates more than 'we did a random survey'. It lets a reader distinguish coverage, selection, response, and measurement. These distinctions also show where an improvement should be made instead of assuming that increasing the final count solves every problem.

8. Compare explanations for a surprising result

Sampling problems often appear when two reports disagree. Suppose an open online poll finds 75 percent support for a new club schedule while a register-based survey finds 40 percent. One explanation is that opinion changed between the two dates. Another is that the question wording differed. A third is that the open poll attracted a different mix of people. Before selecting a favorite story, list the observations each explanation must account for.

Inference to the best explanation compares candidate explanations of the evidence and favors one that explains it well relative to the alternatives considered. It is not a deductive guarantee. In this case, inspect dates, wording, access routes, duplicate handling, and respondent groups. If the reports were collected on the same day with identical wording, a rapid change of opinion and a wording difference become less supported explanations. If the open poll link was posted only in a supporters' chat, selective access becomes a stronger candidate.

Explanatory fit should involve a mechanism. 'The online result is wrong because online polls are bad' merely supplies a label. 'The link reached only the supporters' group, excluding most other members' identifies a route by which the answers could differ. A good explanation connects details of the collection process to the pattern found. It should also respect contrary facts; if many opponents demonstrably answered the open poll, the simplistic all-supporters story needs revision.

Consider what new evidence would distinguish the candidates. A dated archive can check timing. Saved questionnaires can check wording. Distribution records can show who received a link. A follow-up sample drawn from the complete membership register can examine whether the first respondent mix was unusual. These checks are more informative than choosing the explanation that makes your preferred policy look good. Specify the predicted finding before looking, so the hypothesis cannot be adjusted freely to fit every possible result.

The best explanation among those listed may still be inadequate. You might have omitted duplicate submissions, a recording error, or a difference in who counted as a member. Being the best of a weak list does not make an explanation true. Keep a conclusion proportionate: 'Selective access is currently the best-supported explanation of the discrepancy' is different from 'selective access is certainly the only cause'. Further evidence can strengthen, weaken, or replace the account.

This method connects critical reading with practical repair. If distribution records support a coverage problem, widen the invitation route. If nonresponse is the main concern, investigate missing replies. If duplicated responses distorted a count, audit the records. Explaining the discrepancy matters because different explanations call for different actions. Naming bias without specifying its mechanism does not tell a researcher what to fix.

9. Survivors are a selected group

A study of currently open shops asks their owners which practices helped them succeed. That can describe the surviving shops' practices. It cannot, on its own, establish which practices distinguish success from failure because closed shops are missing. If both successful and failed shops used the same practice, its prevalence among survivors would not explain the difference. The relevant comparison needs information about the cases that did not remain available for interview.

This is survivorship selection: entry into the observed group depends on having reached the point at which measurement occurs. It can affect old buildings, functioning devices, continuing club members, and businesses still trading. Always ask what had to happen for a case to remain visible. The missing cases may be especially relevant to the outcome being studied, but their absence does not automatically tell us their characteristics.

Repair begins by defining the original cohort and tracing outcomes, rather than starting only from current survivors. A club studying retention can begin with everyone who joined in January and record who remained by December. Interviewing only December members answers a different question. If some original members cannot be traced, report that missingness rather than silently classifying them as satisfied or unsuccessful. The selection mechanism determines which inference the available evidence can support.

10. Two schedule surveys and a traceable selection rule

A fictional club has 200 registered members. An open poll link posted in a 40-person supporters' chat receives 32 replies, of which 24 favor a new schedule. Its reported support is 24 divided by 32, or 75 percent. A separate uniform draw from the complete register produces 50 completed responses, of which 20 favor it, or 40 percent. Both questionnaires used the same wording on the same day in this exercise.

The arithmetic alone does not decide which report should describe the whole club. The open link's distribution gives only the supporters' chat an invitation route. Even within that chat, eight people did not reply. The register-based draw reaches beyond that subgroup. Under the stipulated complete responses to the selected invitations, it has a better connection to the whole membership question, though ordinary sampling variation remains.

The 35-percentage-point discrepancy is consistent with selective access influencing the open poll. That is a supported explanation because the route is documented and relates directly to the opinion being measured. It is not an accusation that any respondent lied or that every chat member agrees. Nor does it establish that the population support is exactly 40 percent merely because one better-designed sample returned that value.

A careful report gives the two selection procedures, response counts, proportions, and target population. It can use the register-based survey as the more defensible current estimate while checking how much sampling uncertainty remains. The open poll still describes its respondents and may reveal their concerns, but it should not be promoted to a census of the entire club.

11. Where this goes wrong

Calling a subgroup useless. It may describe that subgroup well while failing to justify a wider estimate.

Calling haphazard selection random. Easy access is not a documented chance mechanism.

Assuming every-tenth selection is always safe. A random start and the ordering's relation to the outcome both matter.

Claiming a known direction from self-selection alone. A selection risk does not by itself establish whether a mean is too high or too low.

Treating the best available explanation as certain. Important alternatives may remain untested or unconsidered.

12. Early arrivals do not cover the late bus

  1. State the target question.

    How do all students at this school travel here?

    The population includes students arriving at different times.

  2. Identify the selection rule.

    Ask the first 60 through the gate at eight.

    The rule selects on arrival order rather than by a uniform draw.

  3. Identify an excluded subgroup.

    The late bus arrives at eight twenty.

    Its passengers cannot be among the students already selected at eight.

  4. Connect exclusion to the outcome.

    Bus travel influences arrival time in this case.

    The route into the sample is related to the transport mode being measured.

  5. State the limited result.

    The answers describe early arrivals; a whole-school transport estimate needs a broader design.

    The subgroup's evidence is useful without being representative of every student.

13. A random draw with missing replies

  1. Record the initial selection.

    100 names drawn uniformly from a complete register.

    This provides a documented probability mechanism for invitations.

  2. Record actual response.

    Only 30 selected people answer.

    The realized sample differs from the invitation list.

  3. Calculate missing replies.

    100 − 30 = 70.

    Those selected cases supply no measured answer yet.

  4. Ask the relevant representativeness question.

    Do responders and nonresponders differ on the outcome?

    Selective response could reintroduce a connection between inclusion and answers.

  5. Preserve the distinction.

    Good initial selection; response bias still needs investigation.

    Random invitations do not guarantee that the responding subset retains the same properties.

14. Explain two conflicting schedule results

  1. Calculate the open-poll result.

    24 ÷ 32 × 100 = 75%.

    The percentage uses the actual open-poll respondents.

  2. Calculate the register-sample result.

    20 ÷ 50 × 100 = 40%.

    This percentage concerns a separately selected group.

  3. Inspect two candidate explanations.

    Timing and wording were the same in the stipulated case.

    Those checked details weaken explanations based on different dates or questions.

  4. Inspect access to the open poll.

    Its link appeared only in a supporters' chat.

    The distribution route is connected to the opinion being measured.

  5. Prefer the currently supported mechanism provisionally.

    Selective access is a strong candidate explanation of the discrepancy.

    The documented access difference explains why the respondent mixes may differ.

  6. Keep the conclusion revisable.

    Audit duplicates and records before claiming selection is the sole cause.

    The best available explanation is not guaranteed to exhaust every possible explanation.

15. A periodic production line

  1. State the sequence rule.

    Every tenth item comes from machine B; other items come from machine A.

    The ordering contains a repeating feature that may relate to quality.

  2. Inspect the proposed sample.

    Inspect positions 10, 20, 30, and so on.

    This fixed phase selects only machine B items.

  3. Your turn: work this step out. Its working is at the end of the packet.

    State the coverage problem.

16. Guided practice

A survey is run like this: the first sixty students through the gate at eight in the morning were asked. Is the way the sample was chosen sound or skewed?

17. Guided practice

An open poll gets 24 supporting replies out of 32. A register-based sample gets 20 supporting replies out of 50. Complete the proportions and their difference; the counts alone do not settle the selection explanation.

  1. Calculate open-poll support.

    24 ÷ 32 × 100 = open percent.

    The denominator is the open poll's actual respondents.

  2. Calculate register-sample support.

    20 ÷ 50 × 100 = register percent.

    The second result uses its own sample size.

  3. Compare the two percentages.

    Open percentage minus register percentage = gap percentage points.

    This describes the discrepancy that the selection mechanisms must help interpret.

18. Guided practice

A survey is run like this: sixty names were drawn at random from the full school register. Is the way the sample was chosen sound or skewed?

19. Practice

A district survey uses only a landline list, excluding households without landlines. Among invited listed households, many people do not answer. Connect each fact to its specific sampling issue; neither fact establishes the exact direction of the final error.

This task has no paper form; do it on a device.

20. Practice

Two polls disagree. Saved records show identical wording and collection on the same day. Distribution records show the high-support poll was accessible only through a supporters' chat. Connect these checks to the provisional explanatory judgments they support.

This task has no paper form; do it on a device.

21. Somewhere new

A line alternates in ten-item batches: the tenth item of every batch comes from machine B; the other nine come from A. A sampler always takes positions 10, 20, 30, and so on. Construct the supported links about its coverage.

This task has no paper form; do it on a device.

22. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

23. Test question

Researchers draw 100 names uniformly from a complete membership register. They receive 30 replies; the other selected members do not answer. They have not studied how responders differ from nonresponders. Construct the supported links without assuming either perfect representativeness or a known bias direction.

This task has no paper form; do it on a device.

24. What you can do now

You can state the selection rule, identify coverage and nonresponse risks, and distinguish the initial draw from the realized respondents. Explain why a record restricted to surviving aircraft can omit decision-relevant damage patterns, and which assumptions an armour recommendation would need. Next: what a sample's count actually permits you to say about the whole group.

Working for the steps left to you

15. A periodic production line, step 3

The result cannot alone estimate both machines' overall output quality.

A mechanical every-tenth rule is not automatically a representative sample.