Nolege News

Mathematics

A hospital can have worse survival than its rival while being better in every single department. Both figures are correct.

By ·14 September 2026·9 min read

🌐 इस लेख को हिन्दी में पढ़ें
A hospital can have worse survival than its rival while being better in every single department. Both figures are correct.

In short: Simpson's paradox occurs when an association present in every subgroup reverses on aggregation, because the groups differ in composition and the totals weight them differently. This guide works through the mechanism, the documented admissions and kidney-stone cases, why referral hospitals and difficult cohorts systematically look worse, and the genuinely useful rule: whether to trust the split or the total depends on whether the grouping variable is a confounder or lies on the causal path — a question the data alone cannot answer.

Here is a claim that sounds like a mistake. Hospital A has a lower survival rate than Hospital B overall. Hospital A also has a higher survival rate than Hospital B in cardiac cases, in trauma, in paediatrics, in oncology — in every department there is. No figure has been falsified and no arithmetic is wrong.

This is Simpson's paradox, and the reason it deserves attention is not that it is a curiosity. It is that the situation occurs constantly in real data — in medicine, education, hiring, policy evaluation — and that the two contradictory answers are both genuinely derivable from the same dataset. Deciding which one means something requires information that is not in the table.

How a reversal is built

The mechanism is mundane once seen. It needs two ingredients: subgroups that differ in their baseline outcomes, and subgroups that differ in size between the things being compared.

Take a worked example. A private hospital and a government referral hospital each handle straightforward cases and difficult cases. Straightforward cases survive far more often than difficult ones, whichever hospital treats them — that is the first ingredient.

Suppose the private hospital treats 900 straightforward cases with 99% survival, and 100 difficult cases with 60% survival. Its overall survival is about 95%.

The referral hospital treats 100 straightforward cases at 100% survival — better — and 900 difficult cases at 65% survival — also better. Its overall survival is about 68.5%.

Better in both categories, and far worse in total. Nothing has gone wrong. The overall figure is a weighted average, and the weights are completely different: the private hospital's total is dominated by easy cases and the referral hospital's by hard ones. The aggregate is answering a different question from the one it appears to answer — not "which hospital treats patients better" but "what happens on average to the patients who walk through each door", and those doors receive different people.

This is why referral centres, tertiary hospitals and any institution that accepts the cases nobody else will take are systematically punished by raw outcome tables. It is also why a school that admits weak students can look worse than a school that selects strong ones while teaching better, and why a doctor who takes the hopeless cases has worse numbers than one who declines them.

Two documented cases worth knowing

The best-known real instance comes from graduate admissions at the University of California, Berkeley in 1973. Aggregate figures showed men admitted at a substantially higher rate than women, which looked like clear evidence of bias against women. Examined department by department, the pattern largely disappeared and in several departments reversed — women were admitted at equal or higher rates than men within departments. The explanation was that women had applied disproportionately to departments with low admission rates for everyone, and men to departments that admitted most applicants. The published analysis of this became one of the standard illustrations in statistics.

A second, from medicine, concerns treatments for kidney stones. One treatment showed a better success rate than another for small stones, and a better success rate for large stones, yet a worse success rate overall. The reason was that the more invasive treatment was preferentially used on the large, difficult stones. The treatment was being blamed for the cases it had been chosen for.

Both cases have the same shape: an unevenly distributed third variable — department choice, stone size, case difficulty — that influences both which group you end up in and how you were likely to do anyway.

The paradox is not that the numbers disagree. It is that people expect a total to be a summary of its parts, when a total is a weighted average and the weights carry an argument of their own.

So which number should you believe?

Here is the part that is genuinely useful and usually left out: you cannot resolve this from the data alone. Whether the split figures or the combined figure answers your question depends on what the grouping variable is, causally — and that is knowledge about the world, not about the table.

If the grouping variable is a confounder — something that influences both which group a unit lands in and its outcome, while not being a consequence of the treatment — then the aggregate is contaminated and the subgroup figures are the ones that mean something. Case difficulty is a confounder: it determines which hospital a patient reaches and it determines their prognosis. Comparing hospitals without accounting for it compares patient populations rather than hospitals.

If the grouping variable lies on the causal path — if it is an effect of the thing being studied rather than a cause — then splitting by it is the error, and the aggregate is correct. A blunt example: if a drug works by lowering blood pressure, and you split your results by blood pressure achieved, you have removed the mechanism by which the drug acts and can make a working drug look useless. Adjusting for a mediator does not clean the comparison, it destroys it.

The practical consequence is that "we controlled for everything we had" is not a statement of rigour. Adjustment is only as good as the reasoning about which variables belong in the adjustment, and adding more variables can make an estimate worse rather than better. This is why the discipline of stating the assumed causal structure before analysis has become central to modern statistical practice.

What professionals do about it

The standard tools are worth recognising because they appear everywhere once you know their names.

Standardisation is the everyday fix. Age-standardised mortality rates exist because comparing crude death rates between a young population and an old one measures demography rather than health. The standardised rate asks what each population's mortality would be if both had the same age structure, which makes the comparison about the thing you meant.

Case-mix adjustment does the same for hospital league tables, attempting to account for how sick the incoming patients were. It is necessary and it is also gameable, which is why publishing raw and adjusted figures together, along with the adjustment method, matters more than either number alone.

Randomisation is the reason controlled trials occupy the position they do. If assignment to treatment is random, then the confounders — the known ones and the ones nobody thought of — are distributed evenly across groups by construction, and the aggregate comparison is valid without needing to have identified them. That is the entire logic of the randomised trial, and it is why observational data requires so much more argument to reach the same conclusion.

Reading numbers in ordinary life

A few habits follow directly, and they are not technical.

When you see an aggregate comparison, ask what populations are being compared and whether they are composed alike. When a rate is quoted, ask "per what" — per patient, per attempt, per capita, per rupee — because the denominator often smuggles in the whole disagreement. Be suspicious when a comparison involves an institution that selects its intake, or one that accepts what others refuse. And when a subgroup analysis reverses a headline, do not treat that as automatically the deeper truth either: subgroup analyses on small numbers are also where spurious findings breed.

The useful instinct is not cynicism about statistics. It is the recognition that a number is a summary produced by choices, and that the choices are usually more interesting than the number.

Why it matters for students and researchers

Simpson's paradox is a good entry point into what has become one of the more consequential shifts in applied mathematics — the move from describing associations to reasoning explicitly about causal structure, using formal tools that make assumptions visible rather than implicit. That framework does not create knowledge from nothing; it makes the argument auditable, which is the point.

The practical reach is wide, and for India it is immediate. Comparisons across states, districts and institutions are the basis of a great deal of policy, and populations differing in age structure, urbanisation, baseline health and selection into services are exactly the conditions in which aggregation reverses things. A state that reports higher rates of a disease may simply be diagnosing it more. An institution with poor outcomes may be the one accepting the hardest cases. Work that carefully disaggregates Indian administrative data, and is explicit about the causal assumptions behind any adjustment, is both tractable and genuinely scarce.

That kind of methodological work sits within the scope of Recent Trends in Mathematics (ISSN 3139-6364), a peer-reviewed journal dedicated to emerging developments across the mathematical sciences. For students in mathematics, statistics, economics and the health sciences, the paradox is worth internalising early, because it is the clearest demonstration that a dataset does not contain its own interpretation.

Frequently asked questions

What is Simpson's paradox?

A situation where an association observed within every subgroup of a dataset reverses when the subgroups are combined. It arises when the groups differ both in their baseline outcomes and in their relative sizes across the things being compared.

Is Simpson's paradox a mistake in the calculation?

No. Both the subgroup figures and the aggregate figure are arithmetically correct. The aggregate is a weighted average, and the weights differ between the groups being compared, which is what produces the reversal.

Which number should be trusted, the subgroup or the total?

It depends on the causal role of the grouping variable. If it is a confounder — influencing both group membership and outcome — the subgroup figures are meaningful. If it is an effect of the treatment being studied, splitting by it is an error and the total is correct.

Why do referral hospitals show worse outcomes?

Because they receive a higher proportion of difficult cases. Their overall survival rate reflects the composition of their patients rather than the quality of their care, which is why case-mix adjustment exists.

What was the Berkeley admissions case?

Aggregate 1973 graduate admissions figures appeared to show bias against women, but department-level analysis showed women were admitted at equal or higher rates within departments. Women had applied disproportionately to departments with low admission rates for all applicants.

Does adjusting for more variables always improve an analysis?

No. Adjusting for a variable that lies on the causal path between the treatment and the outcome removes part of the effect being measured and can make a real effect disappear. Which variables to adjust for is a question about causal structure, not a matter of including as many as possible.