Back to the on-screen lesson ·
Random, systematic, stratified and cluster samples stand in for a population only when drawn fairly; intervals and proportional allocations follow from simple ratios.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to design a sample, compute its interval or allocation, and explain why a fair method matters more than a large size.
You have worked with data from maps, rasters and satellites. This lesson turns to data people collect themselves, in surveys and fieldwork, where the question is not only how to measure but which places, people or plots to measure at all.
| Term | What it means |
|---|---|
| Population | Every member the study is about: all households, all streams, all plots. |
| Sample | The members actually measured. |
| Sampling frame | The list or map the sample is drawn from. |
| Systematic sample | Every k-th member, with k the population over the sample size. |
| Stratified sample | A sample drawn from each group in proportion to its size. |
| Bias | A tendency for a sample to differ from the population in one direction. |
A sample stands in for a population only if it is chosen fairly.
Another way: picture
Picture tasting a pot of soup. Stir it first, and one spoonful tells you about the whole pot. Skim only the top, and a thousand spoonfuls still miss the vegetables sitting at the bottom. Sampling design is the stirring.
Another way: steps
A reproducible investigation leaves enough information for another person to regenerate the result. Begin with a question, study population, place, period and decision criterion. Preserve raw observations and a data dictionary with units, missing-value codes and collection methods. Record source versions, CRS, permitted uses and file checksums where available. Work on copies. Log exclusions with reasons, coordinate transformations, join keys, duplicate handling, raster cell size, analysis extent, masks and software versions. A screenshot alone cannot explain which records or settings produced it.
Validation must be designed before seeing the most attractive result. Reserve observations that will not be used for fitting an interpolated surface. Where nearby sites resemble each other, a random split may make prediction seem easier than predicting a new area; consider spatially separated validation sites. Keep sites and dates relevant to the population claim. A convenient roadside sample cannot establish conditions in inaccessible woodland simply because it has many rows.
In an invented soil-moisture investigation, three withheld sites have measured values 10, 14 and 18 percent. A model predicts 11, 12 and 20 percent. Subtract observations from predictions to obtain residuals 1, minus 2 and 2 percentage points. Absolute errors are 1, 2 and 2, so mean absolute error is 5 divided by 3, about 1.67 percentage points. A model can have small mean signed error because positive and negative errors cancel, so inspect individual residuals as well. Do not refit on these sites and still call the same comparison independent validation.
Repeat with a justified alternative setting while holding inputs and validation sites constant. Record whether the estimated priority area changes, and report both runs, their errors and the reason for any preference. If the selected method depends strongly on a parameter, communicate that instability. Reproducibility shows how a result was produced; external validation and uncertainty checks address whether it is useful.
In a simple random sample every member of the population has the same chance of being chosen, as when numbered households are drawn by a random number generator. Chance alone decides, so no one's judgment can tilt the result.
In the field, random points can be placed on a map by drawing random coordinates inside the study area, then navigating to them with GPS.
A systematic sample takes every $k$th member after a random start, where $k$ is the population over the sample size. From $1200$ households, a sample of $60$ takes every twentieth.
Along a line, the same idea gives a transect: points every ten meters from a river's edge up a slope, or every kilometer along a highway. Transects cross gradients, so they show how things change with distance.
A stratified sample splits the population into groups, or strata, and samples each in proportion to its size.
| Group | Households | Share | Sample of 200 |
|---|---|---|---|
| Urban | 6,000 | 60% | 120 |
| Suburban | 3,000 | 30% | 60 |
| Rural | 1,000 | 10% | 20 |
Every group is guaranteed a place, which a random sample might miss by chance.
Visiting households scattered across a state is expensive. A cluster sample picks a few areas, such as census blocks, at random and surveys every household in them.
It saves travel, but neighbors tend to be alike, so a cluster sample needs more people than a random one to reach the same precision.
Field samples come in three shapes. Point samples measure at single spots: a soil core, a water sample. Line samples follow a transect. Area samples use quadrats, squares of set size in which everything is counted.
Each shape can be placed randomly, systematically or by strata.
Bias creeps in through the frame, the method or the response. A frame of landline phone numbers misses households with only cell phones. Sampling only near roads misses what lies far from them. People who answer surveys may differ from those who do not.
More responses do not fix any of these; they only make a biased answer more confident.
A proportional sample of a group that is one percent of the population may hold too few members to say anything about it. Surveys often oversample such groups on purpose, then weight each response by how many people it stands for when adding up the results.
Weighting restores the right proportions for totals while keeping enough cases for the small group.
Checking an answer. Allocations must add up to the whole sample. An interval times the sample should give back the population. A transect has one more point than gaps.
Dividing the population by the sample is allowed for an interval because each pick stands for the same number of members. Multiplying by a group's share is allowed because proportional allocation keeps every member's chance equal.
Adding one point to a transect is allowed because a line with points at both ends has one more point than spaces.
The U.S. Census Bureau counts everyone once a decade, but its American Community Survey samples about three and a half million addresses every year to measure income, commuting, housing and more between censuses. It uses a stratified design and weights, and publishes a margin of error with every figure.
The U.S. Forest Service's Forest Inventory and Analysis program lays a grid of permanent plots across the nation's forests and revisits them, a systematic sample of trees.
A fair sample still carries chance error, which shrinks as the sample grows, but slowly. For a percentage, a rough guide to the margin of error is one over the square root of the sample size. A thousand people gives about $1 / \sqrt{1000} \approx 0.03$, or plus or minus three percentage points, which is why so many national polls survey about a thousand.
Quadrupling the sample only halves the margin. Four thousand people give about plus or minus one and a half points, at four times the cost. So survey designers choose a size that answers the question well enough, and spend the rest of the budget on a better frame and more follow-up with people who did not answer, which reduces bias that no sample size can fix. Good reports print the margin of error next to every figure, so readers know how close a result really is.
The most common slip is dividing the sample by the population for an interval, which gives a fraction. Another is giving strata equal shares instead of proportional ones.
A third is forgetting the starting point on a transect. A fourth is believing a large sample cannot be biased.
Every figure from a survey or field study carries its sampling design with it. The next lesson looks at what fieldwork data can and cannot show once it is collected: spread, outliers and the limits of a small sample.
In 1936 the magazine Literary Digest mailed ballots to about ten million Americans, drawn largely from telephone directories, car registrations and its own subscribers, and got about 2.4 million back. It predicted that Alf Landon would defeat President Franklin Roosevelt.
Roosevelt won in a landslide. In the Depression, people with phones and cars were better off than most and leaned toward Landon, and those who bothered to mail back a ballot differed again from those who did not. George Gallup, polling a far smaller sample chosen to match the country, called the winner.
The episode is still the textbook case: the frame and the method, not the size, decide whether a sample can be trusted.
How much timber, carbon and wildlife habitat do American forests hold? The U.S. Forest Service answers with its Forest Inventory and Analysis program, which places permanent plots on a systematic grid across all land, public and private, and sends crews to measure every tree on the plots that fall in forest.
Because the grid is systematic and the plots are revisited over the years, the program can report not only how much forest there is but how it is changing: growth, harvest, fire and disease, state by state.
The design has limits it states openly. Rare trees and small patches may miss every plot, so figures for small areas carry wide margins of error, and states add denser plots where they need sharper local answers.
It is natural to trust a survey with more responses: surely a hundred thousand answers beat a thousand. But if the method leaves some people out, or lets people choose themselves, adding responses repeats the same tilt more loudly.
Fair methods, such as random, systematic or stratified sampling from a complete frame, make a modest sample trustworthy; unfair ones make a huge sample misleading.
A list has $1200$ households and a survey needs $60$. Divide.
$1200 \div 60 = 20$
Households per pick.
Choose a random start.
$\text{say, the seventh}$
Between one and twenty.
List the first picks.
$7, 27, 47, 67$
Every twentieth.
Name the risk.
$\text{a pattern in the list}$
Every twentieth could be a corner lot.
A county has $6000$ urban, $3000$ suburban and $1000$ rural households. Add them.
$10000$
The population.
Find each share.
$60\%,\ 30\%,\ 10\%$
Of households.
Allocate a sample of $200$.
$120,\ 60,\ 20$
Sample times share.
Check the sum.
$120 + 60 + 20 = 200$
The whole sample.
Name the equal-share slip.
$67 \text{ each}$
Rural households overcounted.
A transect runs $200$ m from the water to the dunes, with points every $10$ m. Count the gaps.
$200 \div 10 = 20$
Spaces.
Add the first point.
$21$
Points at both ends.
Say what to record at each.
$\text{sand size, plant cover, height}$
Change with distance.
Say why a transect suits a beach.
$\text{it crosses the gradient}$
Wet to dry, bare to vegetated.
Add a second transect.
$\text{at a random place along the beach}$
One line may be unusual.
Say how to record the design.
$\text{method, spacing, start}$
So others can judge it.
Find the sampling fraction.
$150 \div 15000 = \frac{1}{100}$
One in a hundred.
Apply it to the stratum.
$3000 \div 100 = 30$
Households.
Check against the whole.
A town has $4000$ households on a list, and a survey needs $100$ of them. For a systematic sample, every how many households should be chosen?
Complete the worked solution: a city has $2500$ renter households and $7500$ owner households, and a survey of $120$ is stratified in proportion. Find the renters' percent of households, the renters sampled and the owners sampled.
Find the renters' percent.
$\frac{\text{renters}}{\text{renters} + \text{owners}} \times \text{a hundred} =$ p
Their share.
Find the renters sampled.
$\text{sample} \times \text{renters' share} =$ i
Proportional.
Find the owners sampled.
$\text{sample} - \text{renters sampled} =$ j
The rest.
Say why stratify here.
$\text{renters and owners differ}$
In moves, costs and views.
Match each sampling method to how it chooses.
| every member has the same chance of being chosen | every k-th member of a list or line is chosen | each group is sampled in proportion to its size | a few whole groups are chosen and everyone in them surveyed | |
|---|---|---|---|---|
| random sampling | ||||
| systematic sampling | ||||
| stratified sampling | ||||
| cluster sampling |
A county has $6000$ urban, $3000$ suburban and $1000$ rural households. Fill in how many of a $200$-household stratified sample come from each group.
| households | |
|---|---|
| urban households sampled | |
| suburban households sampled | |
| rural households sampled |
A survey samples $150$ of $15000$ households, allocated to strata in proportion to their size. Write the sample drawn from a stratum of $x$ households.
Answer:
A class samples vegetation along a transect $200$ m long, at a point every $10$ m, starting at one end and finishing at the other. How many points do they sample?
Answer: points
A school district with $2400$ students in $5$ schools surveys $160$ families about bus routes, in proportion to each school's enrollment. How many families should come from a school of $600$ students?
Answer: families
Mark supported sentences establishing reproducibility, validation and its limit. A fictional stream archive retains raw records, dates, CRS, cleaning rules, transformations, both settings and exclusions. 2 sites are withheld from both candidate interpolation fits. Setting A has smaller prediction errors there; both outputs are retained.
This task has no paper form; do it on a device.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A survey samples $150$ of $15000$ households, allocated to strata in proportion to their size. Write the sample drawn from a stratum of $x$ households.
Answer:
You can reason about sampling design. Explain why the 1936 poll with millions of responses got the election wrong.
26. Your turn: a survey samples $150$ of $15000$ households proportionally. How many come from a stratum of $3000$?, step 3
$15000 \div 100 = 150$
The whole sample.