Back to the on-screen lesson ·
Every report has two groups in it: the one that was measured and the one the sentence is about.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will take any report that argues from evidence and name its two groups: the sample, which is who was actually measured, and the population, which is who the claim is about. You will then find one person or thing in the population that had no way of getting into the sample, which is the fastest test there is for whether a generalization has outrun its evidence.
You can already find the claim in a short argument and say which sentences are offered as reasons for it. That is the whole of what you need here. This lesson adds one question to ask of the reasons: who exactly was measured?
| Term | What it means |
|---|---|
| Target population | The group and period the claim is intended to describe. |
| Realized sample | The cases whose data were actually collected. |
| Sampling frame | The list or access route used to reach possible sample members. |
| Census | Measurement intended to include every member of a defined population. |
| Analogy | An inference from a known case to another using relevant similarities. |
| Disanalogy | A difference that may weaken the proposed transfer between cases. |
Evidence covers a group, and that group is nearly always smaller than the group the claim is about. The whole of this course is the habit of saying both out loud — this is who was measured, this is who the sentence is about — and noticing when they are not the same group.
Take a sentence like students here hate the new schedule. The claim is about students here — that is the population. The evidence behind it might be eleven people complained in the corridor, and eleven people in a corridor is the sample.
Neither group is the villain. A sample is how anybody finds anything out: nobody weighs every bag or asks every student. The move that matters is simply saying both groups out loud, one after the other, in this order:
Once those two lines are written down, the question is that enough? asks itself. Left unwritten, it never comes up at all.
Another way: picture
Draw a circle for the target population and mark the cases actually measured. The diagram identifies coverage. A suitable sampling method can support estimates beyond those cases; it does not promise that every unmeasured person has the same value as every measured person.
Another way: steps
Three questions, in this order, of any report you meet:
A report can involve more than two relevant groups. The target population is the group the question concerns. The sampling frame is the list or route used to reach potential participants. The realized sample is the set whose data were actually recorded. These groups can differ for different reasons. A complete school register may include everyone the question concerns, while only some invited students return a questionnaire. That is different from using a cafeteria line that never includes students eating elsewhere.
Start with the claim and give it a boundary. 'Students at this school this term' is more precise than students. It identifies an institution and a period. Then identify the evidence boundary: perhaps 40 students in one class completed a questionnaire on Tuesday. Finally ask how those students were reached. The class register is a route to those 40 students, but it is not a route to every other class. A student from another class might have had no chance of selection under that method.
Distinguish exclusion from nonresponse. A customer who never submitted a website review is absent from the realized sample of reviews. They were not necessarily unable to write one. A customer without access to the review system may face a coverage barrier. Both absences matter, but for different reasons. Naming the mechanism helps identify the repair: improving access addresses coverage, while following up nonresponders addresses response patterns. Simply collecting more reviews from the same route may address neither.
A sample need not resemble the population in every conceivable respect. It needs an appropriate connection for the question being answered. A height survey and a schedule-preference survey may require attention to different characteristics. Similar shoe colors do little to justify either inference. Age distribution may matter to height; travel arrangements may matter to schedule preferences. Explain why a difference could affect the measured outcome instead of merely listing every way people differ.
The sample can still tell us something valuable even when a wider inference is weak. Forty classmates' recorded answers are evidence about those answers, assuming the collection is accurate. They may identify an issue worth investigating across the school. The problem is claiming a school-wide proportion or universal opinion without the required support. Preserve the useful narrow finding while marking the wider uncertainty. This is more informative than either accepting the headline unchanged or dismissing the study as worthless.
A census measures every member of a defined population. It removes sampling uncertainty for that specified group but does not guarantee perfect measurement. Some responses may be misunderstood, recorded incorrectly, or missing despite the intended census. Nor does a census of one class become a census of a whole school when the headline expands. Population coverage and measurement accuracy are separate questions, and both need attention before a broad conclusion is accepted.
An argument by analogy uses a known case to suggest something about another case that shares relevant features. A library may reason that a return reminder will work for its readers because a similar library used it successfully. The source case is the library with the observed result. The target case is the library considering the change. The proposed conclusion concerns the target, so the argument needs a reason that the similarities matter to the result being transferred.
Write four lines: what happened in the source, what is proposed for the target, which similarities connect them, and which differences could disrupt the connection. If the reminder worked because readers received messages before a predictable due date, then access to messages and the borrowing schedule matter. Having the same wall color does not. The quality of an analogy depends on the relevance of the similarities, not merely the number that can be listed.
Imagine two school libraries with the same borrowing period and the same reminder system. A pilot in the first finds that returned-on-time books increased after reminders began, with a suitable comparison helping assess the change. The second school also has functioning contact details and similar borrowing routines. Those facts give a reason to investigate whether the approach may transfer. They do not guarantee the same numerical improvement, because attendance, response habits, and other conditions may differ.
Now add a material difference: most readers in the second school cannot access messages outside school hours, while reminders arrive in the evening. This disanalogy concerns the mechanism by which the reminder was supposed to help. It is more important than whether the buildings are the same age. A better target plan might send reminders during an accessible period and test whether the result transfers. Identifying a disanalogy can improve an intervention instead of merely defeating discussion.
An analogy is usually defeasible: further relevant differences can weaken its conclusion. Match the wording to that status. 'This gives a reason for a monitored pilot' may be better supported than 'the same result is certain'. A single source case also leaves questions about how typical that case was. Several independently studied source settings with a consistent mechanism can provide broader support, while ten descriptions of the same pilot do not become ten separate studies.
This relates directly to sample and population. Generalizing from observed people to unobserved people depends on which differences matter to the claim. Transferring a policy from one setting to another depends on a similar question about the mechanism. Neither inference requires pretending that the cases are identical. Both require an explicit, relevant bridge and a conclusion whose strength fits the evidence. The next lesson will inspect how the original sample was selected before relying on that bridge.
A database row need not represent a person. One customer can place several orders, one account can log several sessions, and one passenger can board twice. Before reporting how many people were measured, identify what each row represents and whether repeat records were linked. Deduplicating changes a record count into a person count only when the identifiers and linkage are suitable.
A fictional study app has 500 registered accounts. During one week, 200 accounts open the app and produce a total of 8,000 recorded usage minutes. The app publishes 'Our users spend 40 minutes a week studying here.' Dividing 8,000 by 200 gives 40 minutes, but the denominator identifies active accounts, not every registered account.
If the records reliably show that the other 300 accounts recorded zero use during that week, then the mean across all 500 registered accounts is 8,000 divided by 500, or 16 minutes. Both means can be correct because they answer different questions. Forty describes recorded use per active account. Sixteen describes recorded use per registered account, including the zero-use accounts. Neither automatically measures time per person, because one person could have multiple accounts or several people could share an account.
The report should therefore name the unit, period, and denominator. A precise sentence is: 'Among the 200 accounts active this week, recorded app use averaged 40 minutes; across all 500 registered accounts, it averaged 16 minutes.' Calling all recorded time studying would require a further assumption about what users were doing while the app was open.
This example shows why a narrow finding need not be false to support a misleading headline. The total minutes and arithmetic can be accurate while the population label changes. Checking boundaries catches the problem before anyone debates whether the app is useful. The usefulness question may need additional evidence about learning, but identifying active versus registered accounts is the first necessary correction.
Reading the sample as the population. The report says students think; the measurement says forty students in one class said. Nothing was falsified — a wider group was written where a narrower one belongs, and the sentence became a bigger claim on the way.
Treating a sample as worthless. The opposite error, and it is just as wrong. Forty students in one class is real evidence about those forty. It is also decent evidence about the class, and weaker evidence about the year, and the honest report says which.
Thinking a big number fixes it. Two hundred thousand website reviews are still two hundred thousand reviews written by people who decided to write one. Size does not reach a group the method never touched.
State the target of the headline.
All students at the school this semester.
The headline's scope determines the population being claimed about.
State the recorded sample.
40 students in one class answered on Tuesday.
These are the people whose responses the report actually contains.
Identify the route used to reach them.
The teacher asked that class.
This route does not reach students in the other classes.
Name an excluded population member.
A student attending a different class.
That person belongs to the headline's population but had no chance of inclusion by this method.
Write the warranted narrow report.
These 40 respondents gave these answers; school-wide opinion remains to be investigated.
The revision preserves the measured finding without silently expanding its coverage.
Fix the original population.
The 30 students in one class.
The teacher's stated question concerns that class only.
Identify who was measured.
All 30 students were measured accurately.
Under the exercise's assumptions, the measurement covers the full defined group.
Classify the original coverage.
A census of that class.
No class member lies outside the measured group.
Change the proposed conclusion.
The same mean describes the whole year group.
The target population has expanded while the evidence has not.
Identify the new inferential task.
Justify generalization from this class to the other classes.
Complete coverage of a subgroup does not automatically cover the larger group.
Record the unit and time period.
Accounts during one week.
The data count accounts, not necessarily distinct people.
Record the total usage.
8,000 minutes.
The same numerator is used in both comparisons.
Calculate the active-account mean.
8,000 ÷ 200 = 40 minutes.
Only accounts that opened the app belong to this denominator.
Include the reliably recorded zero-use accounts.
200 + 300 = 500 registered accounts.
The expanded population includes the accounts absent from the active group.
Calculate the registered-account mean.
8,000 ÷ 500 = 16 minutes.
Including zero-use accounts changes the denominator and the question.
Report both scopes accurately.
40 per active account; 16 per registered account in this week.
The two correct results are compatible because they describe different populations.
Identify the known source result.
A studied library reminder improved timely returns under its tested conditions.
The source result is the evidence being transferred.
Check a relevant target similarity.
The target uses the same borrowing period and readers can receive the reminders.
Those features bear on the proposed reminder mechanism.
Limit the conclusion and test a difference.
Four pieces of evidence. Match each one to the group it actually covers.
| Students who used the cafeteria that day | Customers who chose to write a review | People who had already joined a running club | Bridges that have not fallen down | |
|---|---|---|---|---|
| Everybody in Tuesday's cafeteria line was asked what they ate | ||||
| The star ratings left on the shop's website were counted | ||||
| The running club timed all of its own members | ||||
| An engineer measured every bridge in the county still standing |
An app records 8,000 minutes across 200 active accounts this week. There are 500 registered accounts; reliable logs show zero use for every inactive account. Complete the population check.
Identify the active-account denominator.
Active accounts measured: active.
This group is smaller than the registered-account population.
Calculate the active-account mean.
Total minutes divided by active accounts = active_mean minutes.
The denominator must match the active-user claim.
Calculate the registered-account mean.
Total minutes divided by registered accounts = all_mean minutes.
Including the zero-use accounts changes the population represented by the mean.
The evidence is this: every student who waited in line at the cafeteria on Tuesday was asked what they ate. Which group does that evidence cover?
A school has 480 students. A survey records answers from all forty students in one class and makes a claim about every student at the school. Give the realized sample size, the target population size, and the number of target students whose answers were not recorded.
Sample size s; population size g; unobserved target students u.
A reminder improved returns in a studied library because readers received it before a shared due date. A target library uses the same borrowing period and verified accessible reminders. Its walls are also the same color. Construct only the supported links: relevant similarity can justify a monitored target pilot, not a guaranteed identical result.
This task has no paper form; do it on a device.
A shop counted the star ratings left on its website last month. For each group, say whether that evidence reaches them.
| The evidence reaches this group | The evidence says nothing about this group | |
|---|---|---|
| Customers who wrote a rating last month | ||
| Customers who bought something last month and said nothing | ||
| People who looked at the shop and did not buy | ||
| Everyone who has ever bought from the shop |
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Counters record every ferry boarding at one terminal during one specified week: 500 boardings. They also record that 80 of those boardings occur before eight in the morning. They do not record people who leave the line before boarding, and they have no other weeks' records. Construct the supported links.
This task has no paper form; do it on a device.
You can name the sample and the population of a report and find somebody the evidence could not have reached. Tell someone why counting more website reviews does not tell you about the customers who wrote none. Next: how the sample was chosen, and how selection affects the strength and scope of an inference.
13. Transfer a library reminder cautiously, step 3
A target pilot is warranted; inaccessible evening messages could prevent the same result.
An analogy supports a conditional expectation rather than a guaranteed identical effect.