Suppose you’re choosing between two tutoring programs. Program A boasts that 82% of its students passed the final exam. Program B reports just 43%. Easy choice?
Not yet. Break the results down by how difficult the students’ courses were, and Program B comes out ahead in both groups. There’s no arithmetic mistake. The reversal has a name—Simpson’s paradox—and it reveals a question worth asking whenever someone gives you a single impressive percentage: Who, exactly, is being counted?

The numbers behind the reversal
Imagine each program taught 100 students. Here are their results:
| Course difficulty | Program A | Program B |
|---|---|---|
| Easier courses | 81 of 90 passed (90%) | 19 of 20 passed (95%) |
| Harder courses | 1 of 10 passed (10%) | 24 of 80 passed (30%) |
| All courses | 82 of 100 passed (82%) | 43 of 100 passed (43%) |
Among students taking easier courses, B’s pass rate is five percentage points higher. Among those taking harder courses, B’s is 20 points higher. Yet combine everyone, and A appears to win by 39 points.
The key is that the programs taught very different groups. Ninety of A’s 100 students took easier courses; only 20 of B’s did. A’s overall figure is dominated by a group likely to pass. B’s is dominated by a group facing a much tougher exam.
An overall rate is a weighted average
A weighted average gives each group an influence proportional to its size. To find Program A’s overall pass rate, multiply each group’s pass rate by its share of A’s students, then add:
A: (90% × 90%) + (10% × 10%) = 82%.
The first 90% in the opening product is the share of A’s students in easier courses; the second is their pass rate. For Program B, the group sizes change:
B: (20% × 95%) + (80% × 30%) = 43%.
That’s the entire trick. We aren’t comparing the two programs using the same mix of students. The difference in weights is large enough to overwhelm B’s higher pass rates within each group.
Equal total enrollment doesn’t fix this: both programs have 100 students. What matters is how those students are distributed between easier and harder courses.
Ask what would happen with the same mix
One way to make a more useful comparison is to give the groups the same weights. Suppose we want to compare the programs for a population split evenly between easier and harder courses. Using each program’s observed pass rates, we get:
A: (50% × 90%) + (50% × 10%) = 50%.
B: (50% × 95%) + (50% × 30%) = 62.5%.
This calculation is called standardization: we apply both sets of rates to one shared population mix. The 62.5% is not a pass rate B actually recorded. It’s a projected rate for that specified 50–50 mix, assuming B’s group-specific rates would hold there.
Neither comparison is automatically the “right” one. If you need to know how many students passed in the groups these programs actually taught, the original totals answer that question. If you’re trying to compare their results for similar mixes of students, the standardized figures are more informative.
Why a fairer comparison is not yet proof
The reversal does not prove that Program B’s teaching caused better outcomes. Students may differ in preparation, attendance, or other ways not captured by “easier” and “harder.” Perhaps particularly motivated students chose B. Those differences could affect pass rates too.
The group sizes also deserve attention. A’s 10% pass rate in harder courses comes from just one pass among 10 students. A few different exam results would move that percentage substantially. Looking inside an overall number is essential, but the numbers inside it still need scrutiny.
To investigate which program causes better results, you’d want comparable students assigned to the programs—ideally at random—and enough students in each group to judge the results with some confidence.
The question to carry with you
Simpson’s paradox can appear wherever an overall rate blends unlike groups: school results, hospital outcomes, sports statistics, or a company’s delivery times. You don’t need to memorize its name to spot it.
When one percentage seems to settle a comparison, ask: What groups are mixed together, and do the groups have the same weights on both sides? If the answer is no, check the rates within comparable groups before declaring a winner. The headline number may be correct—and still answer a different question from the one you meant to ask.


Leave a Reply