clinicians.build · interactive · September 20, 2026 · built on Thompson et al., Health Services Research 2016

The Grade That Doesn’t Keep

Medicare scores every hospital in the country on readmissions and docks up to 3% of its Medicare pay for the result. Line up two measurement windows that share no patients, and 87% of the second grade is not explained by the first.

Primary source: Thompson MP, Kaplan CM, Cao Y, Bazzoli GJ, Waters TM. “Reliability of 30‑Day Readmission Measures Used in the Hospital Readmission Reduction Program.” Health Services Research 2016;51(6):2095–2114 (doi:10.1111/1475-6773.12587, PMID 27766634) — found that readmission rates for medical conditions fall below the reliability benchmark for most hospitals, and that “approximately 25 percent of payments for excess readmissions were tied to unreliable” rates.
Data: CMS Hospital Readmissions Reduction Program (Provider Data Catalog, released Aug 13 2026; 18,330 rows, 3,055 hospitals) via MIMI Labs, joined to the Nov 2023 vintage of the same file. Two non-overlapping three-year windows: Jul 2018–Jun 2021 and Jul 2021–Jun 2024. 7,532 hospital–condition pairs have a published ratio and a published denominator in both.
Read the reliability paper → Or read today’s newsletter →
the number
87%
of a hospital’s 2021–24 excess readmission ratio is not explained by its own 2018–21 ratio.
Same hospital. Same condition. Same CMS model. Different three years. r = 0.354 across 7,532 pairs.

The excess readmission ratio is a verdict. Above 1.00 and your patients came back more than CMS’s model says patients like yours should have; below 1.00 and they came back less. The penalty is capped at 3% of base operating DRG payments — real money, taken on the strength of that number.

So here is the test the number never has to pass. Plot every hospital twice: its ratio in the window ending June 2021, and its ratio in the window that starts the next day. No shared patients. No shared discharges. If the ratio is measuring something the hospital is, the points should sit on the diagonal.

0
hospital–condition pair worst decile of 2018–21, when traced the diagonal: grade unchanged line of best fit
pairs shown
correlation r
grade not carried
crossed 1.00

The cloud is round. The fit line through it has a slope of 0.37 against a diagonal of 1.00 — the signature of regression to the mean, not of improvement. Hospitals that scored badly drifted up toward 1.00. Hospitals that scored well drifted down toward 1.00. And 40% of all pairs crossed the 1.00 line entirely, changing sides on the only question the penalty asks.

The obvious objection, and what happens when you test it

Small hospitals, you’d say. A hospital with sixty heart-failure discharges is going to bounce around, and that bouncing is what’s eating the correlation. It’s the right objection. Drag the slider above and it is partly right — and you can watch exactly how far it gets you.

correlation between the two windows, at each volume floor
Each point recomputes r on only the hospitals clearing that discharge floor in both windows. Points with fewer than 30 surviving pairs are dropped.
the 80/20 lens Raising the floor does raise the correlation — pooled across conditions, from r = 0.354 at every hospital to r = 0.620 at the 128 pairs carrying 1,000+ discharges in both windows. So sample size is real, and it is measurable. But r = 0.62 is r² = 0.38: even among the highest-volume hospital–condition pairs in the country, 62% of the grade still doesn’t carry forward. The noise story explains part of it. It does not rescue the number.

What the funnel-plot intuition gets wrong here

The reflex when you see a rate compared across wildly different sample sizes is to reach for a funnel plot — small units scatter wide, big units cluster tight. Run it on this data and it comes out backwards. Sort heart-failure hospitals into five volume groups and the spread of the ratio grows with size: SD 0.044 in the smallest fifth (median 87 discharges), 0.078 in the largest (median 754).

That’s not because small hospitals are consistent. It’s because CMS already knows they aren’t. The ratio comes out of a hierarchical model that shrinks a low-volume hospital’s estimate toward the national average before publishing it. The bouncing has been pressed flat at the source.

Which makes the headline worse, not better. These numbers have already been smoothed, and they still don’t replicate.

what this is not

Why a clinician already knows this

Medicine built a room for exactly this problem. A group convenes, on a schedule, and asks whether a decision was good — including when the patient walked out fine. The whole premise of that room is that the outcome is not allowed to grade the decision, because outcomes contain luck and decisions don’t.

Then the same profession accepts a number like this one, and treats it as a verdict, because it arrives with a dollar figure attached.

The hospital that came in at 1.14 and is now at 1.04 didn’t fix anything. The hospital at 0.88 that is now 0.94 didn’t break anything. Both of them have a story about why. Only one of those stories got written down, and it’s the one where the number was right.