Back to the on-screen lesson ·

What a count out of a sample permits

Turning a count into a share of the sample, and into the range of values the whole group could still hold.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will turn a count out of a sample into a share of that sample, and then work out the interval of values the whole group could still take — from the count you found, up to that count plus everybody who was never asked. You will be able to say which sentences a sample settles exactly, which it only suggests, and why the gap between the two closes only when the number not asked reaches zero.

2. What you already have

You can name the sample and the population, and judge the rule that picked the sample. Now suppose the rule was a good one. There is still a gap between what the sample shows and what the whole group is like, and this lesson measures it.

3. Words for this lesson

TermWhat it means
Sample proportionThe observed positive count divided by the number of observed cases.
Point estimateOne proposed population value inferred using the sample and relevant assumptions.
Logical boundsThe smallest and largest counts compatible with the confirmed cases and stated information.
Sampling variationDifferences between results that arise from which cases a chance process selects.
PrecisionThe degree of variation or uncertainty in an estimate under its specified design.

4. A count settles the sample and bounds the group

Suppose 14 of 40 students asked say they walk to school. Two different sentences are available, and they are not equally safe.

The first is about the sample: 14 of these 40 walk, which is $35$ percent of them. This is not an estimate. It is a count of the observed reports; treating it as an exact property count assumes accurate responses and recording.

The second is about the population: about 35 percent of the year group walks. This is a guess, and it may be a good one — but it is a different kind of sentence, and the honest report says which it is making.

Between those two there is something you can work out with certainty. In a year group of 200, the 14 walkers are definitely walkers, so the true number is at least $14$. The $160$ students not asked could all walk, so the true number is at most $14 + 160 = 174$. The count of walkers in the year is somewhere in the interval from $14$ to $174$ — and every number in it is still possible.

That interval is embarrassingly wide, and it should be. It is what one sample proves, as opposed to what it suggests.

Another way: picture

Draw a row of 200 boxes. Forty of them are opened: fourteen have a tick and twenty-six are blank. The other hundred and sixty are still shut. The smallest possible total is the fourteen ticks you can see; the largest is those fourteen plus every shut box turning out to be a tick.

Another way: steps

From a sample of $n$ out of a group of $N$, with $c$ saying yes:

  1. Share of the sample: $c \div n$, times a hundred.
  2. Not asked: $N - n$. Write this number down; it is the width of your ignorance.
  3. Lowest possible total: $c$.
  4. Highest possible total: $c + (N - n)$.
  5. Say which sentence you are making — about the sample, or about the group.

5. What a bigger sample does and does not do

Take the same year group of 200 and the same share, $35$ percent, found in samples of different sizes.

AskedSaid yesNot askedCertain range for the year group
2071807 to 187
401416014 to 174
1003510035 to 135
20070070 to 70

The range narrows as the sample grows, and it closes completely only when everybody has been asked. This is the honest arithmetic of a sample, and it is why the last row is not a sample at all.

Notice what changed between this table and the one in lesson 2. There, multiplying the sample by a hundred changed nothing, because the rule was shutting a group out. Here the rule is fine and size is exactly what helps. Both are true, and which one applies depends on the method, not on the number.

6. Three different numerical statements

Keep three answers separate whenever a report gives a sample count. The sample proportion describes the observed group. A population estimate extends beyond it using assumptions about selection and measurement. A logical bound describes every population count still compatible with the confirmed cases alone. These answers can all be useful without being interchangeable. The very wide bound in this lesson is not a claim that a well-designed sample supplies no information about which compatible counts are plausible.

For a group of 24 members, suppose eight are checked and three have the property. The observed proportion is 3 divided by 8, or 37.5 percent. Applying that proportion to 24 gives a point estimate of nine members. The confirmed-case minimum is three. There are 16 unexamined members, so the maximum is 3 + 16 = 19. Nine lies inside the logical bounds, but the bound calculation does not privilege it. A sampling design and statistical reasoning would be needed to justify it as an estimate.

The upper bound can also be calculated by counting confirmed negatives. Of the eight examined members, five lack the property. Those five cannot contribute to the total, so at most 24 − 5 = 19 can have it. This second route checks the first. If the two routes disagree, inspect whether you subtracted the whole sample instead of only confirmed negatives, or added the observed positives twice. A good numerical answer includes a reason for each endpoint.

These are counts of distinct members of one fixed group. They assume the classifications are correct and that the eight examined cases belong to the group. If one person answered twice, the sample size is not eight distinct people. If the group changed between records, the population total needs a time boundary. If the property was mismeasured, the supposedly confirmed counts may need correction. The arithmetic can be valid while the inputs are unreliable; do not treat a formula as a substitute for checking those inputs.

The feasible counts are whole numbers, although an interval conveniently displays their inclusive lower and upper bounds. Saying the count is between three and nineteen does not claim that 7.4 people is a possible count. For this course's interval activities, use the closed endpoints requested and interpret the population count as integer-valued. A proportion or a mean can be fractional; the number of members cannot.

If further evidence establishes additional positives or negatives, update the bounds rather than starting over with a new guess. Finding two more positives raises the minimum from three to five. Finding three more negatives lowers the maximum from nineteen to sixteen. The unknown pool shrinks by five. Bounds are a transparent record of what the evidence excludes, which is why they are useful even before a more advanced estimate is attempted.

7. Size helps precision when the design supports the inference

A larger suitable sample usually reduces the random variation of an estimate because more independent cases contribute. In a simple random sample, an unusually high or low result becomes less common as sample size grows, under the relevant assumptions. This is a statement about repeated sampling behavior, not a promise that every larger realized sample lies closer to the population value than every smaller one. Chance can still produce an unusual collection.

The logical bounds in this lesson behave differently from a statistical uncertainty interval. Their width is exactly the number of distinct unexamined members when no other information constrains those members. They ask what is possible, not what is likely under a sampling model. A statistical interval instead uses a design or model to describe uncertainty in estimation. Do not label the range from three to nineteen a 95 percent confidence interval. No confidence level was calculated to obtain those endpoints.

A large number of records is not always a large number of independent cases. Five hundred readings from one sensor may tell us much about that sensor over time but little about differences among five hundred sensors. Fifty answers from each of ten classrooms may share classroom influences. If the question concerns variation across schools, repeatedly measuring one classroom cannot create coverage of unmeasured schools. Count the units relevant to the inference, not just the rows in a file.

More observations also cannot automatically remove a systematic coverage problem. A survey of 6,000 gym users may describe gym users more precisely than a survey of 60. It still does not directly measure residents who never attend. That does not make the extra observations meaningless; it means they improve a different part of the task. Precision concerns how variable an estimate is under repeated collection. Bias concerns systematic differences introduced by the design or measurement. A result can be very precise about the wrong target.

Increasing sample size has diminishing returns for many ordinary estimates. Under common independent-sampling assumptions, uncertainty often shrinks roughly with the square root of sample size: using four times as many cases can approximately halve a standard-error measure. This is a preview, not a rule to apply blindly to every survey. Clustering, unequal selection, nonresponse, and a large sampled fraction of a finite population can change the calculation. The design must be specified before a numerical margin is claimed.

An honest report therefore names the observed count, sample size, selection procedure, and target group before offering an estimate or uncertainty statement. If it gives only a percentage with many decimal places, ask how much evidence supports that precision. Four yes answers out of ten produce exactly 40 percent of those ten, but printing 40.000 percent does not make a population estimate more informative. Extra digits are not extra observations, and a larger headline number is not a substitute for a sound method.

8. An equipment audit with known and unknown cases

A fictional club has 80 rechargeable lamps in a fixed inventory. An auditor tests 20 distinct lamps using a reliable procedure and finds that nine need replacement batteries. The observed replacement proportion is 9 divided by 20, or 45 percent of the tested lamps. A manager multiplies 45 percent by 80 and obtains an estimate of 36 lamps across the inventory. That estimate is different from what the tested cases alone establish with certainty.

There are 80 − 20 = 60 untested lamps. The nine confirmed cases set a minimum of nine needing batteries. If every untested lamp also needs them, the maximum is 9 + 60 = 69. Equivalently, the 11 confirmed lamps not needing replacement rule out more than 80 − 11 = 69. The inclusive count bounds are therefore nine through 69. The estimate of 36 is compatible with the record, but it is not proved by it.

Now the auditor tests ten additional lamps and finds two more needing batteries and eight not needing them. The combined tested count is 30, with 11 positives and 19 negatives. Fifty lamps remain unknown. The updated minimum is 11 and the maximum is 11 + 50 = 61, also 80 − 19. The bounds narrow because additional distinct cases have been classified.

The club can use these figures to distinguish a guaranteed minimum purchase from a planning estimate and from a full inventory decision. Choosing a sensible purchasing strategy would also involve costs and timing, but the reasoning task comes first: label the estimate as an estimate and keep the confirmed-case bounds visible. No budget decision should depend on pretending that all 80 lamps were tested.

9. Where this goes wrong

Reporting the count as a percentage. 14 of 40 becomes 14 percent, which is a different and much smaller claim. Divide before you write the word percent.

Treating the sample's share as the group's share. $35$ percent of the forty is a measurement; $35$ percent of the two hundred is a prediction. Both may be worth saying and they are not the same sentence.

Believing a big sample proves an exact figure. Even 100 out of 200 leaves a range a hundred wide. Certainty arrives only when the not-asked column reaches zero.

Throwing the sample away. The opposite mistake, and the one this lesson's interval is meant to prevent: the interval has a bottom end, and that bottom end is a real, certain, useful fact about real people.

10. A proportion is about the tested group

  1. Identify the observed positives.

    7 music students play an instrument.

    This count supplies the numerator.

  2. Identify the sample size.

    25 students were asked.

    The denominator is the measured group, not the whole year.

  3. Divide count by sample size.

    7 ÷ 25 = 0.28.

    This expresses the positive share of the sample.

  4. Convert to a percentage.

    0.28 × 100 = 28%.

    Percent expresses the same fraction per hundred.

  5. State the scope and assumption.

    28% of these 25 respondents reported playing, assuming accurate recording.

    The calculation does not by itself describe every student in the year.

11. Use positives and negatives to check both endpoints

  1. Record the group and sample.

    N = 24; n = 8; positives = 3.

    These are distinct members of the same fixed group.

  2. Set the confirmed minimum.

    At least 3 have the property.

    The known positives cannot disappear from the total.

  3. Count unknown cases.

    24 − 8 = 16.

    Only the unexamined members remain free to be positive or negative.

  4. Set the maximum.

    3 + 16 = 19.

    The largest total makes every unknown member positive.

  5. Verify using known negatives.

    8 − 3 = 5 negatives; 24 − 5 = 19. Inclusive bounds [3, 19].

    The second route excludes exactly the cases known not to have the property.

12. Update the bounds after a second audit

  1. Record the first lamp audit.

    80 lamps; 20 tested; 9 need batteries.

    The first audit identifies nine positive and eleven negative cases.

  2. Calculate the original unknown count.

    80 − 20 = 60.

    These lamps contribute the initial uncertainty in the count bounds.

  3. State the original bounds.

    9 through 9 + 60 = 69.

    Unknown lamps can all be negative or all positive under the stated information.

  4. Add the second audit without duplicating cases.

    10 new tests: 2 positives and 8 negatives. Totals: 30 tested, 11 positives.

    The second group is explicitly distinct from the first.

  5. Update the unknown count and endpoints.

    50 unknown; minimum 11; maximum 11 + 50 = 61.

    Additional classifications narrow what remains possible.

  6. Check with total negatives.

    11 + 8 = 19 negatives; 80 − 19 = 61.

    The independent endpoint calculation confirms the updated maximum.

13. A one-bus sample

  1. State the observed fraction.

    One late bus out of one checked: 100% of the sample.

    The percentage describes only the checked case.

  2. Count the unobserved buses in the day's fixed group.

    40 scheduled buses − 1 checked = 39 unknown.

    The other buses' outcomes are not supplied.

  3. Your turn: work this step out. Its working is at the end of the packet.

    State the inclusive count bounds.

14. Guided practice

Of the $20$ people asked, $9$ said they grow something at home. What share of the sample is that? Give a whole number of percent.

Answer:

15. Guided practice

A fixed group has 75 members. Of 25 distinct accurately checked members, eight have a property. No other outcomes are known. Complete the observed percentage and possible population-count limits.

  1. Calculate the sample percentage.

    8 ÷ 25 × 100 = percent percent.

    The denominator is the checked group rather than the whole population.

  2. Count members whose outcomes remain unknown.

    75 − 25 = unknown.

    These are the population members outside the observed sample.

  3. Make every unknown case positive to get the maximum.

    Confirmed positives plus unknown cases = maximum.

    No supplied information rules out a positive outcome for any unobserved member.

16. Guided practice

A group has $280$ people in it. $40$ of them were asked, and $14$ of those said they walk to school. Counting only what is certain, how many people in the whole group could say yes? Give every number that is still possible.

This task has no paper form; do it on a device.

17. Practice

Three surveys, three different sample sizes. Fill in the share for each, as a whole number of percent.

People askedSaid yesShare, in percent
play an instrument257
walk to school4014
used the library last month8028

18. Practice

A group of $125$ people was surveyed by asking $25$ of them. How many people in the group were never asked at all?

Answer:

19. Somewhere new

A forester walks a wood of $400$ trees and inspects $50$ of them, finding the beetle in $28$. Counting only what is certain, how many trees in the wood could have the beetle? Give every number that is still possible.

This task has no paper form; do it on a device.

20. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

21. Test question

A fixed inventory contains 140 devices. Fifty distinct devices are accurately checked; 17 have the specified fault. No other outcomes are known. Give the observed fault percentage, the minimum possible total faulty count, and the maximum possible total faulty count.

Sample percentage percent; minimum faulty count minimum; maximum maximum.

22. What you can do now

You can give a sample's share as a percentage, name how many were not asked, and state the interval the whole group's count must lie in. Tell someone why a sample of one still proves something. Next: what the words some, most and all permit, which is the same question asked of a sentence instead of a survey.

Working for the steps left to you

13. A one-bus sample, step 3

At least 1 and at most 1 + 39 = 40 late buses.

A single verified case gives a real minimum without establishing the day's full rate.