Back to the on-screen lesson ·
Random samples, bias, estimating a population from a sample, and how much samples vary.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
In this lesson you learn how a small random sample can tell you about a whole population. You decide whether a sample is fair or biased, turn a sample proportion into an estimate for the population, and use several samples to see how much estimates vary by chance. You also estimate totals you cannot count, such as the fish in a lake, and meet a famous poll that went wrong because its sample was biased.
You can find a fraction, decimal and percent of a number, and you know that a ratio can be scaled up: if 3 out of every 10 are red, then 30 out of every 100 are red. You can find the mean of a set of numbers and describe how spread out data are. In this lesson you use those skills to answer questions about a large group by looking at only part of it. You learn what makes a sample fair, how to turn a sample result into an estimate for the whole group, and why two fair samples almost never give exactly the same answer.
| Term | What it means |
|---|---|
| Population | The whole group you want to know about, such as every student in a school. |
| Sample | The part of the population you actually ask or measure. |
| Random sample | A sample in which every member of the population has the same chance of being chosen. |
| Biased sample | A sample that favors some members of the population, so it gives a misleading result. |
| Convenience sample | A sample of whoever is easiest to reach; it is usually biased. |
| Inference | A conclusion about a population drawn from a sample. |
| Sample proportion | The fraction of a sample with some feature, used to estimate the fraction in the population. |
| Sampling variability | The natural differences between results from different random samples. |
Often you want to know something about a group too big to ask one by one: every student in your state, every fish in a lake, every light bulb a factory makes. Instead you study a sample, a smaller part of the population, and use it to make an inference about the whole.
The sample only helps if it is like the population. The best way to make that happen is to choose it at random, so that every member has the same chance of being picked. Then no group is favored, and the sample tends to be a small copy of the population. If 14 of 40 randomly chosen students ride the bus, then $\tfrac{14}{40} = 35\%$ of the sample rides, and you estimate that about 35% of the whole school rides too.
Two things can go wrong. A sample can be biased, chosen in a way that favors some members, and then its answer is off in a predictable direction. And even a fair sample is affected by chance, so different random samples give slightly different answers. Bias is fixed by choosing fairly; chance is reduced by taking bigger samples or several of them.
Another way: picture
A grid of 400 small squares standing for a school, about a third of them shaded for bus riders. Forty squares scattered across the grid are circled at random; 14 of the circled squares are shaded. A second picture circles 40 squares in one corner instead, where almost none are shaded.
Another way: story
A cook tastes a spoonful of soup to check the salt. If the pot is stirred first, one spoonful tells her about the whole pot. If she skims a spoonful from the top of an unstirred pot, it might taste of nothing but broth. Stirring is what random sampling does.
To choose a random sample, start with a list of the whole population, such as a school roster. Give every name a number, then use a random number generator, or draw numbered slips from a well-mixed bag, to pick the sample. Every student has the same chance, and nobody decides who gets in.
Random does not mean careless or whoever you happen to meet. Asking the first 30 people who walk past you feels random, but the people who walk past at 8 a.m. outside a gym are not like the whole town. A true random sample needs a chance process that gives everyone the same opportunity.
Ask two questions about any sample. Who could not be chosen? And who was more likely to be chosen? If the answer to either is a group that might answer differently, the sample is biased.
A biased sample does not become fair by getting bigger. Asking 5,000 people at a library still gives too high a reading rate. Size cannot fix the way a sample was chosen.
Once you have a random sample, there are two equal ways to estimate a count in the population.
Both are proportional reasoning: the sample and the population are assumed to have the same ratio. Always call the answer an estimate, such as "about 270", because another sample would give a slightly different number.
Suppose four classes each take a random sample of 25 students from the same school and ask whether they sleep at least 8 hours. They might get 11, 14, 9 and 12 yes answers, or 44%, 56%, 36% and 48%. Nobody made a mistake. Each sample caught a slightly different mix of students just by chance. This is called sampling variability.
Looking at several samples tells you how much to trust one. Here the sample percents spread from 36% to 56%, so a single sample of 25 could easily be off by 10 percentage points. Larger samples vary less: samples of 100 from the same school usually land much closer together. Averaging several samples also helps, because the highs and lows partly cancel.
Some populations cannot be listed at all, like the fish in a lake or the beans in a huge jar. Scientists estimate such totals by marking some members and then taking a random sample. If the marked members make up a certain fraction of the sample, they should make up about the same fraction of the whole population.
For example, suppose 50 beans are painted and mixed back into a jar. A random scoop of 60 beans contains 3 painted ones, which is $\tfrac{3}{60} = \tfrac{1}{20}$. If 50 beans are one twentieth of the jar, the jar holds about $50 \times 20 = 1{,}000$ beans. The idea works only if the painted beans are mixed in well, which is the same as saying the sample is random.
A bigger random sample usually gives a better estimate, because chance has less room to push the result around. Think of flipping a coin: 10 flips can easily give 7 heads, which is 70%, but 1,000 flips will almost always land between 45% and 55% heads. Samples of people work the same way.
The surprise is how small a good sample can be. What matters most is the size of the sample, not the size of the population. A well-chosen random sample of about 1,000 adults can estimate the opinion of a whole country almost as well as it estimates the opinion of one city. Going from 1,000 to 4,000 people only makes the typical error about half as big, so pollsters rarely pay for much bigger samples. For a class project, a random sample of 30 to 50 students is usually enough to see the big picture of a school.
Use these steps whenever you reason from a sample.
How to check. Make sure the estimate is less than the population. Check that the population's percent matches the sample's percent. If you estimated several groups, such as favorite lunches, check that the estimates add up to the whole population.
In 1936 the magazine Literary Digest mailed survey cards to about 10 million people and got back about 2.4 million, one of the largest polls ever taken. It predicted that Alf Landon would beat President Franklin Roosevelt. Instead, Roosevelt won about 61% of the vote and 46 of the 48 states. The names came largely from telephone books and car registrations, and in the Great Depression people with phones and cars were wealthier than most voters. Only people who chose to mail the card back were counted, too. George Gallup, using a far smaller sample chosen more carefully, predicted Roosevelt's win. A biased sample of 2.4 million lost to a well-chosen sample of about 50,000.
Factories cannot test every product, because testing can destroy it: you can only crash-test a car or drop-test a phone once. So they test a random sample. Suppose a bakery makes 12,000 granola bars a day and weighs 80 of them chosen at random. If 4 of the 80 are underweight, that is $4 \div 80 = 5\%$, so the bakery estimates about $0.05 \times 12{,}000 = 600$ underweight bars that day and adjusts its machines. If the next day's sample has only 1 underweight bar in 80, the estimate falls to 150, and the fix seems to be working.
The most common mistake is thinking a bigger sample fixes bias. A huge online poll is still made only of people who chose to answer.
The second is treating an estimate as exact. "About 270 students" is right; "exactly 270" claims more than a sample can tell you.
The third is thinking that a sample that differs from another must be wrong. Random samples differ by chance; that is sampling variability, not an error.
The fourth is reporting the sample count as the population count. If 12 of 40 sampled students walk, that does not mean 12 students in the school walk.
The last is calling any unplanned choice random. Asking friends or whoever is nearby is a convenience sample, not a random one.
Write the sample proportion as a fraction.
$\dfrac{21}{60}$
In a random sample of 60 students, 21 ride the bus.
Simplify the fraction.
$\dfrac{21}{60} = \dfrac{7}{20}$
Divide the top and the bottom by 3.
Change it to a decimal.
$\dfrac{7}{20} = 0.35$
Thirty-five hundredths, or 35%.
Apply the proportion to the population.
$0.35 \times 1{,}200 = 420$
A random sample should have about the same share as the school.
Check by scaling the sample up.
$1{,}200 \div 60 = 20, \quad 21 \times 20 = 420$
Both methods give about 420 bus riders.
Find the percent in the library survey.
$\dfrac{45}{50} = 0.90 = 90\%$
Forty-five of 50 adults asked at the library read every week.
Find the percent in the random phone survey.
$\dfrac{26}{50} = 0.52 = 52\%$
Twenty-six of 50 adults picked at random from the town list read every week.
Estimate from the library survey.
$0.90 \times 8{,}000 = 7{,}200$
This is what the library sample would claim.
Estimate from the random survey.
$0.52 \times 8{,}000 = 4{,}160$
This is what the random sample suggests.
Find the gap between the two estimates.
$7{,}200 - 4{,}160 = 3{,}040$
The two samples are the same size but disagree by more than 3,000 people.
Decide which estimate to trust.
$\text{library} \to \text{biased toward readers}$
Everyone at a library is more likely to read, so trust the random survey: about 4,160.
Change each count to a percent.
$\tfrac{11}{25}, \tfrac{14}{25}, \tfrac{9}{25}, \tfrac{12}{25} = 44\%, 56\%, 36\%, 48\%$
Each count is out of 25, and $\tfrac{1}{25} = 4\%$.
Add the four percents.
$44 + 56 + 36 + 48 = 184$
The first step toward their mean.
Find the mean percent.
$184 \div 4 = 46\%$
Averaging the samples gives a steadier estimate than any one of them.
Estimate for the school of 600.
$0.46 \times 600 = 276$
Apply the mean percent to the whole population.
Find the lowest single-sample estimate.
$0.36 \times 600 = 216$
The sample with 36% gives the smallest estimate.
Find the highest single-sample estimate.
$0.56 \times 600 = 336$
The sample with 56% gives the largest estimate.
Describe the variability.
$336 - 216 = 120$
One sample of 25 could be off by dozens of students, so the average of all four, about 276, is the best estimate.
Find the sample proportion.
$\dfrac{9}{150} = 0.06$
Nine out of 150 is 6 out of 100.
Apply it to the day's production.
$0.06 \times 5{,}000 = 300$
The tested phones were a random sample of the day's phones.
Check by scaling up.
A principal wants to know how many of the school's $770$ students ride a bike to school, and will ask $77$ of them. Which sample is most likely to represent the whole school?
Three random samples of 20 students each were asked whether they play a musical instrument. The numbers who said yes were $3$, $7$ and $8$. The school has $400$ students. Complete the worked solution to estimate how many of them play an instrument.
Add the three sample counts.
$3 + 7 + 8 =$ t
Putting the samples together uses all the information you have.
Divide the total by the number of samples.
$\text{total} \div 3 =$ e
The mean count is a steadier estimate for a group of 20 than any one sample.
Multiply by the number of groups of 20 in the school.
$400 \div 20 = 20, \quad 20 \times \text{mean} =$ s
A random sample stands for the population, so each group of 20 students should look about the same.
In a random sample of $20$ students, $4$ walk to school. The school has $220$ students. How many times as large as the sample is the school, and about how many students walk?
k times as large; about answer students walk.
A random sample of $50$ adults was asked whether they bike to work; $4$ said yes. Write that share as a decimal and as a percent.
Decimal: d. Percent: answer%.
Four random samples of students at a school of $1500$ were asked whether they own a pet. The lowest sample percent was $34\%$ and the highest was $40\%$. Between what two numbers of students would the school's estimates fall?
Between lo and hi students.
A town has $12000$ residents. At the town pool, $44$ of 50 people asked said they swim every week. In a random sample of 50 residents, $24$ said they swim every week. Estimate the number of weekly swimmers from each sample.
Pool sample: about p. Random sample: about answer.
Biologists catch $60$ fish in a lake, tag them and let them go. A week later they catch $48$ fish at random, and $8$ of them have tags. What fraction of the second catch is tagged, and about how many fish live in the lake?
Fraction tagged: f. Fish in the lake: about answer.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A random sample of 40 students at a school of $760$ chose their favorite lunch: $16$ chose pizza, $10$ chose tacos and $14$ chose salad. Estimate how many students in the whole school would choose each lunch.
Pizza: about p. Tacos: about t. Salad: about answer.
You can reason from samples. Without looking: why is asking the first 50 people at a gym a poor way to estimate how often a town exercises, and if 18 of 60 randomly chosen students play a sport, about how many of 1,500 students do?
18. Your turn: 9 of 150 randomly tested phones fail a drop test, out of 5,000 made, step 3
$5{,}000 \div 150 \approx 33.3, \quad 9 \times 33.3 \approx 300$
Both ways give about 300 phones that would fail.