Nature Medicine says stop grading the algorithm and start grading whether the patient is better off. American hospitals have been running that experiment in public for fifteen years. Here is the answer, for 2,141 of them.
On Friday, Novartis reported that pelacarsen lowered Lp(a) exactly as designed in 8,323 patients and did not reduce heart attacks, strokes, or cardiovascular death. On Monday, Kristina Lång made the same argument about medical AI in Nature Medicine: the first generation was judged on whether algorithms could match clinicians; the next should be judged on whether human–AI systems improve patients.
The uncomfortable part is that we already have a fifteen-year natural experiment on exactly this question, and it is free. CMS publishes, for every US hospital, a set of process measures — did you complete the sepsis bundle, did you vaccinate your staff, did you co-prescribe opioids and benzodiazepines at discharge — and a set of outcome measures: did the patients die.
Process measures are the surrogate endpoint of health services. They are the AUROC of a hospital. They are cheap to move, easy to audit, and they are what quality departments are actually staffed to improve.
So: if a hospital gets better at the measured step, do fewer of its patients die?
Five process measures. Five outcome measures. Every cell below is the share of the variation in that outcome that is statistically explained by that process measure, across the same 2,141 hospitals. Click any cell to load it into the scatter underneath.
Each dot is one hospital. Size is Medicare admission volume. Pick the axes; the line is the ordinary least-squares fit through all visible dots.
Drag Min volume to the right. Every step throws away small hospitals, and by about 3,000 admissions you are down to roughly 150 of the largest academic centres in the country. Watch what happens to r.
On the sepsis bundle against hospital-wide mortality, the correlation across all 2,141 hospitals is +0.003 — nothing at all. Restrict it to the 155 hospitals with at least 3,000 Medicare admissions and it climbs to +0.127, which, taken literally, says that hospitals better at the sepsis bundle kill more patients. Staff flu vaccination drifts the opposite way over the same slider, from −0.166 to −0.249 by the time 643 hospitals are left.
Neither of those numbers is a discovery. They are what noise looks like when you let someone choose the sample. A slider that quietly shrinks n from 2,141 to 24 will hand you a publishable-looking effect in either direction, on demand. If your model’s eval report contains a subgroup where performance is unexpectedly strong, the first question is not why — it is how many.
The honest version of a null result includes the reasons it might be wrong. Four of them here are load-bearing:
What survives all four caveats is the part that matters to a builder: every one of these process measures was expensive to instrument, is reported quarterly, has a dashboard, and has a person whose job depends on it — and the public data cannot show you what any of them bought.
Your differential generator hits 92% top-3 accuracy. That is a process measure. It is your SEP-1.
The pairs on this page are what happens when an entire industry measures the step it can see for fifteen years without instrumenting the step after it. The recommendation gets generated — then does the attending open it, does it survive the nurse’s read-back, does the order get placed, or does it die in a soft-stop nobody has looked at since 2019?
Every one of those is measurable in weeks, not years, and none of them are on anyone’s model card. Every system can quote its AUROC. Almost none can quote their advice-to-action conversion rate.