Back to the on-screen lesson ·

The table with four boxes in it

Two yes-or-no questions make four kinds of case, and a claim that quotes one of them has compared nothing.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

You will take any claim that one thing goes with another, name the two yes-or-no questions behind it, and lay the evidence out as a table of four boxes with its row totals, column totals and grand total. You will read a named box rather than a total, fill a table in from partial information, and check the grand total twice. You will also be able to name the four boxes as true positive, false positive, false negative and true negative.

2. What you already have

You can say what a sample covers and what a count out of it permits. Every question so far has been about one thing at a time. This lesson is the first about two things at once — a treatment and a result, an alarm and a fire — and two yes-or-no questions make four cases, not two.

3. Reading the table

TermWhat it means
CellCases satisfying one row condition and one column condition together.
Marginal totalA row or column total that combines the other classification's outcomes.
ConjunctionThe event that both specified conditions hold.
Conditional proportionThe share with one property among a stated subgroup.
Conjunction fallacyJudging both A and B more probable than A under the same evidence and inclusive reading of A.

4. Two questions, four mutually exclusive cells

A two-by-two table classifies the same cases using two binary questions. Suppose a diary records 56 days and asks whether a charm was carried and whether the day went well. The definitions of carried and went well must be fixed before counting. Every recorded day must receive one answer to each question; otherwise an unknown category or a clear exclusion rule is needed.

Day went wellDay did not go wellTotal
Carried charm18725
Did not carry charm22931
Total401656

The eighteen good days with the charm establish that the two events occurred together eighteen times in this record. They do not, on their own, establish that carrying the charm improved the chance of a good day. That stronger claim requires a comparison. The observed good-day share is 18/25 = 72 percent when the charm was carried and 22/31, approximately 71 percent, when it was not. These recorded proportions are close.

The table does not prove that the charm has exactly zero causal effect. The difference could reflect chance, other differences between the days, the way good was measured, or some mixture of influences. Its immediate lesson is that quoting eighteen successes hides the denominator and the comparison group. Descriptive similarity and causal equivalence are different claims, just as an observed difference does not establish a causal explanation by itself.

Each interior cell describes a conjunction: both its row condition and its column condition hold. A row total includes both possible column outcomes, while a column total includes both possible row outcomes. The grand total counts every case once. Do not add a row total to a cell inside that row: those cases would be counted twice. When calculating a rate, write the population after the word among before choosing the denominator.

Another way: steps

Name the observational unit and both binary questions. Label rows and columns. Enter known cells. Subtract only when a total and all other cells in that line are known. Check the grand total using rows and columns. For a rate, name the group in its denominator before dividing.

5. Known cells, missing cells, and impossible tables

A table has four interior counts, two row totals, two column totals, and one grand total. These nine displayed numbers are connected by addition, so they are not nine independent facts. If a row totals 25 and its left cell is 18, its right cell is seven. If the left column totals 40, its lower cell is 22. If the grand total is 56, the remaining cell is nine. The completed table can then be checked by adding both rows and both columns separately.

Partial information is not always sufficient to fill every cell. Suppose a study has 100 cases, 60 in row A, and 50 in column B. Let the overlap of A and B be x. The other cells must be 60 minus x, 50 minus x, and x minus ten. All four must be nonnegative, so x may be any integer from ten through fifty. The totals leave many tables possible. Choosing x = 30 without further evidence invents a fact rather than reconstructing one.

The nonnegative-count check can also reveal inconsistent inputs. If a row total is 20 but one cell in that row is 24, no nonnegative value can complete the row. Pause to inspect transcription, category definitions, and whether the total covers the same time period. A negative answer is evidence that the stated inputs or your placement of them cannot all be right. Do not hide it by replacing the result with zero.

Even a perfectly adding table can have poor inputs. A log containing only occasions when an alarm sounded cannot reveal how often fires occurred while it was silent. It needs observations of both alarm states and an independent way to establish whether a fire occurred. Missing cases do not become negative cases merely because they were not recorded. Arithmetic consistency checks the relationships among the numbers; it does not establish completeness or measurement accuracy.

The observational unit must remain stable. If rows count people but columns count visits, a person attending several times may occupy more than one apparent cell. Decide whether the study concerns people, visits, days, or devices, and classify that same unit twice. If a person's status changes over time, specify the observation window. These decisions explain why a four-cell table is appropriate instead of treating the format as suitable for every collection of numbers.

6. Conjunction cannot exceed its component

A conjunction requires both A and B. Every case satisfying A and B also satisfies A, and also satisfies B. Therefore the count in their shared cell cannot exceed either corresponding total. For a randomly selected member of the same fixed group, this becomes the probability rule P(A and B) is no greater than P(A), and no greater than P(B). The rule follows from inclusion, not from any assumption that the properties are independent.

Imagine a fictional archive containing 100 complete records of adult volunteers. Thirty volunteers are engineers. Twelve of those engineers also play the violin. The number who are engineers and violinists is twelve, while the number who are engineers is thirty. If a record is selected uniformly, the probabilities are 12 percent and 30 percent respectively. The more detailed description picks out a subset; adding a requirement cannot enlarge the set of cases satisfying it.

The conjunction fallacy occurs when someone judges the conjunction more probable than one of its components under the same evidence and interpretation. A vivid description of an inventive, musical volunteer can make engineer and violinist sound representative. That feeling does not overturn set inclusion. After conditioning both probabilities on the same description, the conjunction still cannot be more probable than engineer alone. The numerical probabilities may change, but the subset relation remains.

This diagnosis requires a fair comparison. In ordinary conversation, engineer alone might be interpreted as engineer who is not a violinist. Under that reading the options are two disjoint cells, and neither cell must be larger. State that engineer includes engineers who play violin and engineers who do not. Likewise compare both claims using the same evidence. Adding extra information to only one side changes the task rather than violating the conjunction rule.

Do not confuse conjunction with conditional probability. Among engineers, twelve of thirty play violin, so the conditional share is 40 percent. Among all volunteers, twelve of one hundred satisfy both conditions, so the joint share is 12 percent. The different denominators explain why the conditional percentage can exceed the overall engineer percentage. That comparison is not a conjunction fallacy: it compares different mathematical quantities. Write the conditioning group explicitly before judging a probability statement inconsistent.

The same discipline helps with persuasive descriptions. More detail can make a story easier to imagine while imposing more conditions for it to be true. That does not mean detailed stories are always unlikely or misleading; evidence can strongly support each detail. It means plausibility should not be assessed by vividness alone. Ask whether the claim has silently changed from one broad event to a narrower event inside it, and preserve the inclusion relation when assigning probabilities.

7. Test outcomes and comparison rates

For a test table, positive and negative describe what the test reports. True and false describe whether that report agrees with the independently established state. A true positive is a positive result when the condition is present; a false positive is a positive result when it is absent. A false negative misses a present condition, and a true negative correctly reports its absence. The labels do not depend on whether the state itself is desirable.

For a revision guide, 24 passes among 30 users is 80 percent. Twelve passes among 20 nonusers is 60 percent. Comparing 24 with twelve alone mixes the difference in group sizes with the difference in observed success shares. The rate comparison shows a twenty-percentage-point association in these data. It does not settle whether choosing the guide, prior preparation, access to tutoring, or another influence explains that association. The table organizes evidence for further reasoning; it does not supply random assignment after the fact.

8. Auditing a fictional delivery forecast

A distribution club evaluates 100 scheduled delivery days. Before each day, a model forecasts either a delay or no delay, using one fixed definition of delay. Afterwards an independent log records whether a delay occurred. There are 18 forecast-delay days with an actual delay, twelve forecast-delay days without one, six no-delay forecasts followed by a delay, and 64 no-delay forecasts without a delay. Each scheduled day appears exactly once.

The forecast-delay row totals thirty. The no-delay row totals seventy. The actual-delay column totals 24, and the no-delay column totals 76. Both routes give the same grand total of one hundred. Of the thirty warning days, eighteen had a delay: 60 percent. Of the 24 actual-delay days, eighteen had been forecast: 75 percent. These different percentages answer different questions, despite sharing the same numerator.

A manager says a day is more likely to have both a delay forecast and an actual delay than to have an actual delay. For uniform selection from this record, that claim compares eighteen joint cases with 24 actual-delay cases. The shared cell is contained inside the actual-delay column, so its probability is 18 percent rather than 24 percent. The conjunction cannot be larger.

This audit describes recorded performance. It does not guarantee that the next hundred days will have the same delay rate or that the forecast causes delays. It also assumes the independent outcome log is complete. If the model's alerts were the only occasions investigated, the six missed delays could disappear from the record and the apparent performance would be misleading. The table makes that missing-data question visible.

9. Checks that prevent misreading

A cell is not a row total. A recorded zero is not an unobserved outcome. A correct total is not proof of complete data. Similar rates do not prove zero causal effect, and different rates do not prove causation. A conjunction uses both conditions and cannot exceed either component when the evidence and denominator are held fixed.

10. Recover a missing cell

  1. Identify the observed unit and full group.

    50 students, classified by guide use and pass result.

    Both questions classify the same students.

  2. Complete the guide-user row.

    30 users: 24 pass, 30 − 24 = 6 do not.

    The two outcomes exhaust this row.

  3. Complete the nonuser row.

    20 nonusers: 12 pass, 20 − 12 = 8 do not.

    The subtraction uses this row's own total.

  4. Check the columns.

    36 pass and 14 do not; 36 + 14 = 50.

    An independent route verifies the grand total.

  5. Compare observed shares.

    24/30 = 80%; 12/20 = 60%.

    The association is descriptive and does not identify its cause.

11. A conjunction within a marginal total

  1. Fix the reference group.

    100 complete volunteer records, equally likely to be selected.

    Both probabilities must use this same group.

  2. Read the broad count.

    30 are engineers, including violinists.

    The inclusive description contains every engineer.

  3. Read the conjunction.

    12 are engineers and violinists.

    Both conditions pick out a subset of the engineers.

  4. Convert using the shared denominator.

    P(engineer) = 30%; P(engineer and violinist) = 12%.

    The joint event cannot exceed its containing event.

  5. Contrast a conditional question.

    Among engineers, violinists are 12/30 = 40%.

    Changing the denominator creates a different question, not a counterexample to inclusion.

12. Audit the delivery model's two percentages

  1. Place the four joint outcomes.

    Warning: 18 delayed, 12 not; no warning: 6 delayed, 64 not.

    Each day has one prediction and one observed outcome.

  2. Add the prediction rows.

    30 warnings and 70 no-warning days.

    These totals classify days by the forecast.

  3. Add the outcome columns.

    24 delayed and 76 not delayed.

    These totals classify days by what actually happened.

  4. Check the full count both ways.

    30 + 70 = 100; 24 + 76 = 100.

    Agreement is necessary for arithmetic consistency.

  5. Ask how often a warning corresponded to a delay.

    18/30 = 60%.

    Among warning days makes all thirty warnings the denominator.

  6. Ask how often actual delays were warned about.

    18/24 = 75%.

    Among delayed days changes the denominator to the twenty-four actual delays.

13. Sixty seed trays with two compost types

  1. Complete the new-compost row.

    35 trays: 28 sprouted, 7 did not.

    Each of these trays has one binary sprouting result.

  2. Complete the old-compost row.

    25 trays: 15 sprouted, 10 did not.

    The missing cell is its own row total minus sprouted cases.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Check and compare.

14. Guided practice

Here is a study where a workshop logged every time the smoke alarm sounded. The four counts are filled in. Add the row totals, the column totals and the grand total.

There was a fireThere was no fireTotal
Alarm sounded134
Alarm silent2427
Total

15. Guided practice

Across 80 delivery days, there are 12 warned delays, eight warnings without delay, four unwarned delays, and 56 days with neither. Complete the requested totals and warning precision.

  1. Add all warning outcomes.

    Warnings: 12 + 8 = warnings.

    The warning row includes correct and incorrect warnings.

  2. Add all actual delays.

    Delays: 12 + 4 = delays.

    The outcome column includes warned and unwarned delays.

  3. Calculate the delayed share among warnings.

    Warned delays divided by all warnings = precision percent.

    Among warnings specifies the conditioning group.

16. Guided practice

In a study where a workshop logged every time the smoke alarm sounded, there were $80$ cases altogether. $32$ were “Alarm sounded, There was a fire”, $15$ were “Alarm sounded, There was no fire”, and $17$ were “Alarm silent, There was a fire”. How many were “Alarm silent, There was no fire”?

Answer:

17. Practice

A complete register has 100 volunteers: 40 are runners, and 15 of those runners are also cyclists. A volunteer is selected uniformly. Runner includes those who cycle and those who do not. Give the percentage chance of runner, of runner and cyclist, and of cyclist among runners.

Runner broad%; both joint%; cyclist among runners conditional%.

18. Practice

A study where a weather club checked its forecasts against what happened reports its totals and only one of its four counts. Rebuild the rest of the table.

It rainedIt stayed dryTotal
Rain forecast2124
No rain forecast15
Total261339

19. Somewhere new

A weather club checked its forecasts against what happened. On $23$ days it forecast rain and it rained; on $5$ days it forecast rain and it stayed dry; on $6$ days it forecast no rain and it rained anyway; on $38$ days it forecast no rain and it stayed dry. Build the table.

It rainedIt stayed dryTotal
Rain forecast235
No rain forecast638
Total

20. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

21. Test question

A complete log has 120 days. A warning coincides with a delay on 24 days; there is a warning without delay on 16 days; and a delay without warning on six days. Give the count with neither warning nor delay, the count of all delays, and the percentage of all days with both warning and delay.

Neither neither; all delays delays; joint percentage joint%.

22. What you can do now

You can build a two-by-two table, fill it in from totals, check it two ways and read the box a question actually asks about. Tell someone why eighteen lucky days with a charm is not evidence that the charm works. Next: the same four boxes when the condition is rare, which is where good tests start misleading people.

Working for the steps left to you

13. Sixty seed trays with two compost types, step 3

43 sprouted + 17 not = 60; 28/35 = 80%, 15/25 = 60%.

Shares allow a descriptive comparison without treating it as proof of causation.