Back to the on-screen lesson ·

What an accuracy figure leaves out

One percentage stands in for four boxes, and which two it came from decides whether it means anything.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will take a single figure quoted about a test — accurate, reliable, right nine times out of ten — and ask the question that unpacks it: out of whom? You will work out what a test that always answers no would score on the same condition, use that as the score to beat, and name the two kinds of mistake separately, saying which one a given situation cares about.

2. One figure, two places it is quoted

One figure, two places it is quoted

A supplier sells one detector and quotes one figure for it: 97 percent accurate.

In the warehouse

The warehouse uses it to look for smoke. A fire it misses burns the building down.

An alarm on a day with no fire costs an hour of everybody's time and a call-out fee.

At the gate

The same detector is used at a gate to decide whose bag is searched on the way out.

A miss lets one item through. An alarm on somebody carrying nothing has them stopped and searched in front of the line.

The figure itself

The supplier's sheet does not say how many of the tested cases had something to find, or how the 3 percent of errors were split between the two kinds.

3. What you already have

You can fill in the screening table and work out what share of positive results are real. This lesson goes the other way: somebody hands you one percentage, and you work out which of the four boxes it came from — and which ones it hides.

4. Performance measures

TermWhat it means
AccuracyCorrect classifications divided by all evaluated cases.
SensitivityDetected actual positives divided by all actual positives.
SpecificityCorrectly negative results divided by all actual negatives.
Positive predictive valueTrue positives divided by all positive results.
BaselineA specified comparison procedure evaluated for the same task and population.

5. One advertised figure, two decisions

Read the fictional supplier sheet and the two settings. The supplier quotes 97 percent accuracy but does not identify the evaluation population or split its errors into false positives and false negatives. The warehouse and gate descriptions illustrate different consequences of mistakes. They do not provide complete numerical costs, acceptable error limits, or evidence that the detector performs identically at both sites.

The question is what the figure establishes and what a buyer still needs. A missed fire and an unnecessary evacuation differ in consequences. A missed item and an unjustified search differ too. However, the short descriptions do not justify ranking the two buyers' anger or choosing a detector on their behalf. A reasoned response should identify the missing evidence and connect each missing quantity to a decision it could affect.

6. Accuracy combines correct outcomes

Overall accuracy is the number of true positives plus true negatives, divided by the total number of evaluated cases. In a complete record of one hundred messages, a filter that correctly flags 29 spam messages and correctly leaves seventy legitimate messages has 99 percent accuracy. It misses one spam message. A different filter can achieve the same accuracy by correctly flagging all thirty spam messages while incorrectly flagging one legitimate message. The shared percentage hides which error occurred.

An always-negative classifier supplies a useful baseline. If ten of ten thousand components have a defect, saying no defect for every component produces 9,990 correct answers and ten errors, or 99.9 percent accuracy. Its sensitivity is zero: it detects no defects. Its specificity is one hundred percent: it does not falsely flag any nondefective components. Its positive predictive value is undefined because it produces no positive results at all; zero positives cannot serve as a nonzero denominator.

A useful detector need not exceed this baseline's overall accuracy. If missing a defect is much more costly than an additional inspection, accepting more false flags can be worthwhile. That is a decision about consequences and constraints, not an arithmetic fact supplied by the accuracy figure. The baseline exposes how easily a high score can arise in an imbalanced population. It does not establish a universal requirement that any lower-scoring detector be discarded.

Conversely, overall accuracy is not meaningless. It measures the fraction classified correctly on the stated evaluation set, treating each mistake equally. When that is the relevant objective and the evaluation is representative, it can be informative. The error is demanding that this one measure also answer every question about detection, positive predictions, individual cases, and practical value. Give the figure its proper scope rather than treating either a high number or criticism of that number as a complete evaluation.

Another way: steps

State the measured population and time period. Identify the numerator and denominator of each percentage. Recover the four counts when possible. Compare a relevant baseline on the same cases. Examine the two kinds of mistake separately, then connect their consequences to the stated decision.

7. Five related fractions with distinct denominators

The four cell counts determine several useful rates. Sensitivity is true positives divided by all actual positives: TP/(TP + FN). It answers what share of existing cases the test detects. Specificity is true negatives divided by all actual negatives: TN/(TN + FP). It answers what share of noncases the test correctly leaves negative. The false-positive rate is FP/(FP + TN), so it is one minus specificity when that denominator is nonzero.

Positive predictive value is TP/(TP + FP). Its reference group is everyone with a positive result. It answers how many positive results correspond to an actual case in this population. Accuracy is (TP + TN)/(TP + FP + FN + TN), using every evaluated case. The same numerator may appear in sensitivity and positive predictive value, but their denominators differ. A high sensitivity does not, by itself, imply a high positive predictive value.

These quantities are related rather than unrelated. Let prevalence be the actual-positive share of the evaluation population. Accuracy equals prevalence times sensitivity plus one minus prevalence times specificity. It is a weighted average of the two correct-classification rates. If sensitivity and specificity remain fixed while prevalence changes, the overall accuracy can change. Positive predictive value also depends on prevalence, as the previous lesson showed using tables and Bayes' rule.

Naming a denominator precisely prevents ambiguous phrases from concealing the comparison. The phrase false-alarm rate sometimes means false alarms among noncases and sometimes false alarms among all alarms. Those are different rates. State the count and denominator rather than assuming the phrase has one universally understood meaning. In this lesson false-positive rate always uses the noncase group, while the false share among positive results is one minus positive predictive value.

Some rates cannot be calculated in a particular evaluation. If there are no actual positive cases, sensitivity has a zero denominator. If there are no positive predictions, positive predictive value has a zero denominator. Report such a rate as undefined for that dataset, not as zero or one hundred percent. A rate requires a relevant reference group as well as a numerator. Collecting more of the wrong kind of case does not solve the missing-denominator problem.

8. Equal accuracy can hide a decision-relevant difference

Consider one thousand fictional components, one hundred defective. Detector A catches fifty defects, misses fifty, falsely flags no nondefective components, and correctly leaves nine hundred unflagged. Its accuracy is (50 + 900)/1,000 = 95 percent. Detector B catches all one hundred defects, misses none, falsely flags one hundred nondefective components, and correctly leaves eight hundred unflagged. Its accuracy is 90 percent.

Detector A has the higher accuracy, but choosing between them requires the objective. Under an illustrative rule that assigns cost ten to a missed defect and cost one to an unnecessary inspection, A has total error cost 500 and B has total error cost one hundred. Under a different rule assigning cost one to a miss and ten to an unnecessary inspection, A costs fifty and B costs one thousand. Neither detector is simply better for every purpose. The stated costs, population, and error counts make the comparison assessable.

These toy costs do not prescribe policy for factories, medicine, courts, or security. They demonstrate why a choice cannot follow from accuracy alone when mistakes differ in consequence. Some constraints also resist simple monetary addition: a decision may require a maximum miss rate, a minimum inspection capacity, or a prohibition on particular actions. State those requirements explicitly instead of hiding them inside the word reliable.

Threshold changes can alter the balance between misses and false flags. A system that flags weaker signals may catch more actual cases and also flag more noncases. Whether the resulting tradeoff occurs, and how large it is, must be measured for the system. It cannot be inferred from the generic label smoke alarm or from an analogy with another institution. The practical request is for performance at the operating threshold that will actually be used.

The baseline should match the task too. An always-negative rule is natural when the positive state is rare, but an always-positive rule may score better when positives dominate. Established procedures, another model, and human review can provide additional comparisons. Evaluate them on the same relevant cases where possible. Comparing a new detector on difficult cases with an old detector on easy cases can make a numerical difference look like an improvement in ability when the populations caused it.

9. What an evaluation report must identify

An accuracy claim needs a definition of correct and an independent way to establish the actual state. If a detector's own output determines which cases receive verification, its apparent performance may omit missed cases. Ask whether outcome assessment covered positives and negatives, how uncertain labels were handled, and whether the observation period allowed the target outcome to become known. A complete-looking percentage can conceal selective verification.

The size and composition of the evaluation matter. Ninety-nine correct classifications out of one hundred and 9,900 out of ten thousand both produce 99 percent, but they supply different amounts of evidence about future performance. Neither percentage guarantees its own stability. An evaluation from one environment may fail to represent another with different equipment, weather, users, or underlying prevalence. Report the population and conditions alongside the metric so readers can judge relevance.

Aggregate performance can also hide subgroup differences. Suppose a detector makes most errors under poor lighting, while the full evaluation contains mostly bright scenes. Its total accuracy can be high even though the setting of interest is poorly served. Request the relevant subgroup counts, while checking that those groups have enough observations and that the subgroup definitions were not chosen only after noticing favorable results. More detailed reporting should improve understanding rather than create new opportunities to select impressive numbers.

When responding to an advertisement, distinguish what is established, what is not determined, and what would help. Established: the stated proportion of evaluated cases was classified correctly, assuming the report is accurate. Not determined: the error split, performance in another population, and the value of a particular action. Helpful additions: the complete table, evaluation design, operating threshold, and a clear account of the decision's consequences. This is a constructive critique because each requested fact answers a specific question.

10. A structured reply to the supplier

The fictional supplier quotes 97 percent accuracy for one detector in two settings. A careful reply begins with the strongest warranted reading: three percent of cases in the supplier's evaluation were classified incorrectly, assuming accuracy has its stated overall meaning and the report is correct. That establishes neither how many were missed cases nor how many were false flags. It also does not identify the evaluation's base rate or whether it resembles either proposed site.

For the warehouse, request the number of actual fire events in the evaluation, the number detected, and the number of false alarms among nonfire observations. Ask whether the evaluation conditions resemble the warehouse and which operating threshold generated the numbers. The scenario gives a serious consequence for a missed fire, so sensitivity matters, but evacuation costs and observation coverage still matter too. No single number makes the other questions disappear.

At the gate, request the analogous counts for the relevant prohibited item and the frequency and consequences of unjustified searches. The fact that one device is sold for both tasks is not evidence that the same accuracy transfers between them. Detecting smoke and detecting an item are different outcome definitions, and the supplier must support each claimed use with appropriate evidence.

The response should conclude that the advertised percentage is insufficient for a justified choice between operational policies. It should not conclude that the device is certainly useless or that one buyer must accept more errors than the other. A useful critique turns each missing fact into a targeted request and keeps the eventual decision conditional on the evidence and the agreed consequences.

11. Avoid replacing one overclaim with another

Accuracy is neither the probability that a positive is correct nor a useless statistic. An always-negative baseline exposes class imbalance but does not determine which detector serves a costly-error task best. Sensitivity, specificity, prevalence, and predictive value are connected by the same table; they are distinct, not unrelated. A finite trial's perfect detection rate does not guarantee perfect future detection. A decision needs consequences and constraints as well as performance measurements.

12. Calculate the always-negative baseline

  1. State the evaluation population.

    10,000 components, ten with the specified defect.

    The baseline depends on the same population as the proposed comparison.

  2. Apply the constant prediction.

    Every component is labeled nondefective.

    The rule ignores signals and always predicts the more common class.

  3. Count correct and incorrect labels.

    9,990 true negatives; ten false negatives.

    Only components without the defect receive a correct label.

  4. Compute the overall accuracy.

    9,990/10,000 = 99.9%.

    Accuracy weights every evaluated case equally.

  5. State the limitation.

    Sensitivity 0%; positive predictive value undefined.

    The rule detects no defects and produces no positive group to evaluate.

13. One percent error, two different tables

  1. Fix the reference mailbox.

    100 messages: thirty spam and seventy legitimate.

    Both filters are compared on the same class composition.

  2. Construct one 99-percent-accurate filter.

    29 true flags, one missed spam, zero false flags, seventy true negatives.

    It makes one error by missing spam.

  3. Construct another with equal accuracy.

    Thirty true flags, zero misses, one false flag, 69 true negatives.

    It makes one error by flagging a legitimate message.

  4. Verify the shared score.

    29 + 70 = 99; thirty + 69 = 99 correct.

    Each score divides by the same one hundred messages.

  5. Identify what the score hides.

    The first misses spam; the second misfiles a legitimate message.

    Preference requires a view about the consequences of these different errors.

14. Compare detectors under an explicit cost rule

  1. Fix a shared evaluation group.

    1,000 components, one hundred defective.

    Matching the population prevents a prevalence difference from driving the comparison.

  2. Record detector A's errors.

    Fifty misses and zero false flags; 95% accuracy.

    Its higher score comes with incomplete defect detection.

  3. Record detector B's errors.

    Zero misses and one hundred false flags; 90% accuracy.

    Its lower score accompanies complete detection in this record.

  4. Apply the stated illustrative costs.

    Miss costs ten; false flag costs one.

    Consequences determine how errors are weighted for this particular task.

  5. Calculate both total costs.

    A: 50 × 10 = 500. B: 100 × 1 = 100.

    B has lower error cost under this rule despite lower accuracy.

  6. Test the dependence on the decision rule.

    If miss costs one and false flag ten, A costs fifty and B costs 1,000.

    Reversing the costs reverses the preference, so no unconditional ranking follows.

15. A car-fault baseline

  1. Identify actual negatives.

    10,000 cars minus 800 faulty cars = 9,200 nonfaulty cars.

    An always-negative rule is correct exactly on these cases.

  2. Compute its accuracy.

    9,200/10,000 = 92%.

    The baseline reflects the class composition.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Interpret a detector with the same accuracy.

16. Guided practice

A fictional detector advertisement reports 93 percent overall accuracy on its evaluation set. What does that alone establish about the percentage of its positive results that are correct?

17. Guided practice

Detector A makes eight misses and two false flags. Detector B makes two misses and twelve false flags. For this fictional comparison, a miss costs ten units and a false flag costs one. Complete the total costs and their difference.

  1. Apply the costs to detector A.

    Eight times ten plus two gives a_cost units.

    The two mistake types receive their stated weights.

  2. Apply the identical rule to detector B.

    Two times ten plus twelve gives b_cost units.

    A fair comparison uses the same cost rule for both systems.

  3. Compare the totals.

    A's cost minus B's cost is difference units.

    The difference expresses the comparison under this rule, not an unconditional ranking.

18. Guided practice

Among ten thousand schoolchildren, $10$ have a rare pollen allergy. A machine is built that ignores everybody and prints “negative” every time. Out of the $10000$ people, how many of its answers are right?

Answer:

19. Practice

A complete evaluation contains forty true positives, ten false negatives, twenty false positives, and thirty true negatives. Give accuracy, sensitivity, and the positive predictive value as a fraction.

Accuracy accuracy%; sensitivity sensitivity%; positive predictive value predictive.

20. Practice

Construct the support map for a critique of a supplier's accuracy-only advertisement. Use each supplied evidential link once.

This task has no paper form; do it on a device.

21. Somewhere new

A town has exactly twenty rainy and 180 dry days in a complete 200-day record. An always-dry forecast is compared with a forecast that detects fifteen rainy days and incorrectly forecasts rain on ten dry days. Give the always-dry accuracy percentage, the second forecast's accuracy percentage, and its number of missed rainy days.

Always-dry accuracy baseline%; second accuracy accuracy%; missed rainy days misses.

22. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

23. Test question

In a complete 200-component evaluation, a detector has thirty true positives, ten false negatives, twenty false positives, and 140 true negatives. Give overall accuracy and sensitivity as percentages. Under a fictional rule costing a miss five units and a false flag one, also give total error cost.

Accuracy accuracy%; sensitivity sensitivity%; total cost cost units.

24. What you can do now

You can say which group an accuracy figure is a share of, work out the score a do-nothing test would get, and tell a false positive from a false negative. Tell someone why a machine that never looks at anybody can be 99.9 percent accurate. Next: if-then claims, and the one thing an if-then does not say.

Working for the steps left to you

15. A car-fault baseline, step 3

A 92% detector ties this baseline on overall accuracy; its error split remains unknown.

Equal accuracy does not establish equal sensitivity, practical value, or errors on the same cars.