Back to the on-screen lesson ·
separate a reliable method from one fortunate result
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
You will separate a reliable method from one fortunate result, explaining the inference and its limits in a supplied case.
A true answer can result from luck. To assess a method, compare its outputs with an independently specified reference across relevant cases.
| Term | What it means |
|---|---|
| Reliability | A process's tendency to produce accurate beliefs under relevant conditions. |
| Observed accuracy | Correct outputs divided by all evaluated outputs in the sample. |
| Calibration | Correspondence between stated confidence and observed frequencies across suitable cases. |
| Reference class | The class of cases used to evaluate the process. |
Two people identify a covered token correctly. One uses a calibrated imaging device; the other guesses. Their answers match, but their methods differ. The single success does not reveal whether either method tends to produce accurate beliefs. Reliability concerns that tendency under relevant conditions. A method can be reliable without succeeding every time, and an unreliable method can occasionally produce a true answer. This is why the history and structure of a process matter in addition to the truth of its current output.
Imagine a device checked on twenty known samples. It classifies eighteen correctly and two incorrectly. The observed accuracy is eighteen divided by twenty, or ninety percent. That fraction describes performance on those trials. It does not immediately establish ninety-percent accuracy in every environment or a ninety-percent probability that every particular future answer is correct. To move beyond the sample, we need assumptions about how the samples were selected, whether conditions remain similar, and what kinds of errors occurred.
An epistemological reliabilist gives belief-forming reliability an important role in explaining justification or knowledge. Different reliabilist theories formulate that role differently. The useful introductory contrast is with accounts emphasizing reasons a subject can consciously articulate. A child might accurately recognize a familiar voice without being able to explain the acoustic mechanism. A reliability-based approach can regard the quality of that process as epistemically relevant even when the child cannot provide a technical defense of it.
This does not mean that any method with a good-looking score automatically produces knowledge. Gettier cases, environmental luck, relevant defeaters, and the choice of process description all create further questions. Our task is narrower: identify the process, interpret its performance evidence, and limit the conclusion to the conditions that evidence supports. Numerical summaries help make the argument explicit, but the philosophical issue is what those summaries are evidence about.
Another way: Choose the comparison class carefully
A reliability claim always concerns a process operating across some class of cases. 'Reading this instrument in a controlled laboratory' and 'reading this instrument outdoors in heavy rain' may have different records. Combining the results can conceal an important condition. Suppose an instrument is correct on all ten indoor trials and on six of ten outdoor trials. Its overall sixteen-out-of-twenty score is eighty percent, but that average should not erase the difference between the two settings.
The relevant class should match the intended use. If tomorrow's task is outdoors, the outdoor evidence is more directly relevant than the combined average, though its small sample remains a limitation. If tomorrow's setting is unknown, state that uncertainty. Do not choose the class after seeing which result makes your favored method look best. The class should be defended by features of the belief-forming situation that bear on accuracy, such as lighting, distance, instrument condition, or kind of target.
The same episode can be described at different levels. A person may be 'using vision', 'reading a digital display', or 'reading this frozen digital display'. These descriptions collect different comparison cases. This is related to the generality problem for reliabilism: which process type should determine the reliability relevant to the particular belief? An extremely broad description can hide a local defect, while an artificially narrow description can make any one successful episode look perfectly reliable.
A useful classroom response is to state the process and environment before calculating. For example, 'classification by device R on the supplied dry samples under the same lighting'. That does not solve every philosophical issue, but it exposes the assumption for inspection. Another reader can now ask whether wet samples belong in the intended use or whether the lighting condition has changed. Reliability becomes a claim about a specified route, not a compliment attached vaguely to a person or machine.
Another way: Accuracy, calibration and error patterns
Accuracy is the proportion of correct outputs under the chosen scoring rule. Calibration is a different idea: stated confidence levels should correspond appropriately to observed frequencies across comparable judgments. A forecaster who says 'seventy percent' on many occasions is well calibrated in that group if the predicted event occurs about seventy percent of the time, subject to sampling uncertainty. A perfectly confident forecaster with frequent mistakes has a calibration problem even if many individual predictions happen to be correct.
You do not need advanced statistics to see why error patterns matter. A screening procedure might label every sample negative and achieve ninety-five-percent overall accuracy when ninety-five percent of samples are negative. That score conceals its failure to identify any positive sample. In a philosophical discussion, the example shows that a favorable overall number may not support the specific reliability claim being made. Ask which kinds of correct and incorrect answer are being counted and which matter to the inquiry.
Small samples also limit conclusions. Two correct answers out of two give an observed score of one hundred percent, but they provide less information about a stable tendency than a large, varied set of checks. Do not turn this into the opposite dogma that a large sample automatically settles everything. A thousand repeated trials on one easy kind of sample can still leave the difficult cases untested. Size and representativeness are different properties; both can matter to the inference.
Reliability evidence should be gathered independently of the answers being evaluated. If a procedure defines the correct answer as whatever it outputs, its success rate becomes circular. Use a reference whose authority has a separate basis in the task, and disclose disagreements rather than counting them away. The classroom cases stipulate reference labels to keep arithmetic and philosophical interpretation separate. In real inquiry, establishing those labels is itself an evidential task requiring methods, documentation, and sometimes revision.
Another way: What reliability can and cannot establish
A reliable process can produce a false belief on an unusual occasion. This does not erase its entire track record, though the error may reveal a condition requiring a narrower description. Likewise, a process with a poor track record can produce a true belief by luck. The Gettier lesson explains why a successful result may fail to be knowledge when its connection to the truth is accidental. Reliability directs us toward a pattern, but a pattern is not a substitute for examining the current case's known defects.
Suppose a meter is usually accurate, but the user has just received a specific warning that its battery failure makes today's reading unstable. Citing the annual accuracy rate does not answer that warning. The general record and the local defeater must be considered together. This is a common structure in evidence evaluation: background reliability supports reliance until relevant information changes the assessment. Responsible revision attends to the defect actually identified rather than either ignoring it or concluding that all measurements are worthless.
Reliability should also be distinguished from loyalty, consistency, and agreement. A loyal friend can be sincerely inaccurate. A device that always prints the same number is highly consistent but may rarely be right. Several devices can agree because they share one faulty input. Each of these properties may have practical importance, but none is identical to the tendency to produce true beliefs under the relevant conditions. Define the property before deciding what observation measures it.
When writing an analysis, give the observed fraction, name the evaluated process and setting, and state the limit on extension. For example, 'The scanner was correct on eighteen of twenty dry samples, so its observed accuracy here was ninety percent; these trials do not directly establish performance on wet samples.' This statement is informative without pretending that a finite test proves an exceptionless law. It lets a reader distinguish the arithmetic result from the additional inductive claim.
Finally, ask what a new test should change. If the uncertainty concerns wet conditions, repeat tests with appropriately varied wet samples. If the uncertainty concerns independent checking, use a separate reference route. If the concern is that all observations came from one day, vary the date while controlling other relevant features. A test is most useful when its design addresses the specific gap in the reliability argument. Simply accumulating more of the same easy successes can leave that gap untouched.
Another way: Record failures as well as successes
A reliability record must include inconvenient outputs. If a user saves only correct classifications and deletes the rest, the surviving archive cannot provide an honest denominator. Before interpreting a score, check whether every eligible trial was retained under a rule fixed in advance. This is a question about evidence selection, not merely arithmetic. Even a perfectly calculated percentage can mislead when the underlying record systematically excludes failures.
A small archive compares two scanners that identify handwritten shelf codes. In a supplied trial, scanner A correctly reads eighteen of twenty clean labels and six of ten faded labels. Scanner B correctly reads sixteen of twenty clean labels and nine of ten faded labels. All labels have been independently checked against the accession register, so the exercise stipulates the reference answers rather than deriving them from either scanner.
For clean labels, A's observed accuracy is ninety percent and B's is eighty percent. For faded labels, A's is sixty percent and B's is ninety percent. Across all thirty labels, A has twenty-four correct results, or eighty percent, while B has twenty-five, or approximately eighty-three percent. The combined score modestly favors B, but the subgroup pattern tells the archivist something more useful: the observed advantage depends strongly on the condition of the labels.
If the next task involves only faded labels, the faded-label trials are directly relevant. They support preferring B in that setting, while the small trial size and limited variety remain reasons for additional checks. If the next collection has a very different balance of clean and faded labels, the original combined percentage should not be transferred unchanged. Its value partly reflects the original mixture of cases.
The archive can improve the evidence by testing a larger varied set of faded labels, keeping the reference process independent, and recording particular error types such as confusing letters with numbers. A scanner that makes a correct reading once has not thereby earned unrestricted trust. Conversely, discovering one error does not make every output useless. The practical goal is a documented account of where the method works, where it fails, and what human verification is appropriate for the next use.
A hundred-percent score on two easy cases does not establish infallibility. A ninety-five-percent overall score can conceal failure on every rare positive case. Report the tested class and inspect the error pattern before extending a number to a different use. Reliability concerns accuracy, not mere repetition, loyalty, or confidence.
Name the evaluated process.
Device R classifies dry samples under fixed lighting
The conditions specify which reliability claim is being tested.
Record the supplied trial counts.
18 correct out of 20
The reference labels are stipulated independently of R's outputs.
Form the accuracy fraction.
18/20
Correct outputs are divided by all evaluated outputs.
Convert the fraction to a percentage.
90 percent
Multiplying the fraction by one hundred expresses the same proportion.
Bound the interpretation.
Observed accuracy on these dry samples
The calculation alone does not establish performance in every environment.
Read the indoor record.
10 correct out of 10
This subset concerns the first specified environment.
Read the outdoor record.
6 correct out of 10
This subset concerns the environment of the planned next use.
Calculate the combined record.
16 correct out of 20: 80 percent
Combine numerators and denominators rather than ignoring a subgroup.
Calculate the directly relevant subgroup.
Outdoor observed accuracy: 60 percent
The outdoor trials better match the proposed outdoor task.
State the transfer limit.
The combined average conceals environmental variation
The aggregate does not erase the different performance conditions.
Describe the sample composition.
95 negative cases and 5 positive cases
The base distribution determines how an always-negative rule will score.
State the procedure's output.
Negative for every case
The procedure makes no distinction between the two actual categories.
Count correct negative outputs.
95
Every actually negative case receives the matching output.
Count correct positive outputs.
0
The procedure never produces a positive output.
Calculate overall accuracy.
95 percent
Ninety-five of the hundred total outputs are correct.
Evaluate the narrower claim.
This result does not support reliable detection of positive cases
The overall score hides complete failure on the category at issue.
The stated wet-condition trials include twelve correct and eight incorrect outputs.
20 total outputs
Correct and incorrect trials together form the denominator.
Form the observed correct fraction.
12/20
The numerator counts only outputs matching the independent reference.
Express the bounded score.
A process produces 16 correct and 4 incorrect outputs in its stated reference class. Calculate its observed accuracy.
| Response | |
|---|---|
| Right | |
| Total | |
| Percent |
A stated test contains fifteen correct and five incorrect outputs. Complete the observed-accuracy calculation.
Count all evaluated outputs.
Total = total
Both successes and failures belong in the denominator.
Compute the proportion that matched the reference.
15 divided by the total
Accuracy compares correct outputs with every evaluated output.
Convert the proportion to a percentage.
percent percent
The percentage describes the supplied trial and does not guarantee a future result.
A device has 18 correct of 20 dry trials. In the wet subset it has 9 correct and 6 incorrect outputs. Calculate only the wet subset's observed accuracy.
Observed counts (correct outputs / all outputs; do not reduce): right/total. Observed accuracy as a whole-number percent: percent.
A classifier always outputs negative. A reference set contains 18 negative and 2 positive cases. Calculate its overall observed accuracy; do not infer success at identifying positives.
Observed counts (correct outputs / all outputs; do not reduce): right/total. Observed accuracy as a whole-number percent: percent.
A device is tested on 12 bright cases, all correct, and 10 dim cases, of which 4 are wrong. Give dim correct outputs, dim total outputs and dim accuracy as a whole-number percentage. Use original counts.
| Your reconstruction | |
|---|---|
| Dim correct | |
| Dim total | |
| Dim percent |
An archive checks a scanner on faded labels. Of 25 independently checked labels, 5 are misread and the rest are correct. Calculate the observed accuracy for this faded-label trial.
Observed counts (correct outputs / all outputs; do not reduce): right/total. Observed accuracy as a whole-number percent: percent.
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
A method has 27 correct and 3 incorrect outputs in bright conditions. In dim conditions it has 14 correct and 6 incorrect outputs. Calculate only its dim-condition observed accuracy.
Observed counts (correct outputs / all outputs; do not reduce): right/total. Observed accuracy as a whole-number percent: percent.
Without rereading, explain how to separate a reliable method from one fortunate result. Give a fresh case and identify what would change your analysis.
10. Finish a weather-specific audit, step 3
60 percent on these wet-condition trials
A score must retain the setting to which its evidence directly applies.