clinicians.build · interactive · 10 sep 2026

The Supervisor Already Exists

ARPA-H funded the heart-failure agents — and gave Stanford a separate award to build the thing that watches them. Before anyone builds that, look at the supervisor America has been running on heart failure since 2012, and where it bends.

Primary source: ARPA-H, “ARPA-H to revolutionize cardiovascular disease management with clinical agentic AI,” September 2026
Data: CMS Hospital Readmissions Reduction Program file, July 2021–June 2024 discharges, released May 2026, via MIMI Labs

Everyone is building the clinical agent. Almost nobody is building the thing that watches it. This week the federal government split those into two line items and funded them both — heart-failure agents that can prescribe, and a separate Stanford team whose only job is to sit above those agents and catch them drifting.

That second job sounds new. It isn’t. CMS has graded every hospital in America on 30-day heart-failure readmissions since 2012, with money attached. It is the largest supervisory system ever built for a single clinical condition, and its record is the best available evidence about what happens when you try.

The hard part was never measuring. The hard part is that the measurement means something different depending on how much of it you have.

Follow all 2,246 hospitals through the supervisor

Every hospital CMS grades on heart failure, flowing left to right through three questions. How much data did the supervisor have? (discharge volume). What did it conclude? (excess readmission ratio above or below 1.0). And then the one nobody asks: does that conclusion survive an ordinary significance test?

Fourteen years of grading heart failure, in one picture.
graded worse than expected graded better than expected not distinguishable from its own expected rate · ribbon width = number of hospitals
Graded worse
1,149
ratio above 1.0
Of those, confirmed
204
17.8% survive the test
Indistinguishable
1,889
84.1% of all hospitals
Contradicted
0
CMS never points the wrong way

The third column is the whole story

Look at where the ribbons go. 1,149 hospitals are graded worse than expected. 945 of them — 82.2% — cannot be told apart from their own CMS-expected rate by a plain binomial test. Only 204 are confirmed worse. On the other side, 1,097 hospitals are graded better and only 153 are confirmed better. In total 1,889 of 2,246 hospitals — 84.1% — are carrying a public grade, and money, on a number that is not statistically separable from “average.”

One honest note in CMS’s favour, visible as an absence: zero hospitals are contradicted. Not one hospital graded better is confirmed worse by the raw signal, and not one graded worse is confirmed better. The model is not inventing the direction. It is just very quiet about it.

Now look at the left-hand column

The volume bands do not feed the two grades evenly. 67.4% of hospitals with fewer than 100 heart-failure discharges flow into “graded worse.” For hospitals with 800 or more, it is 41.2%. The band means march the same way: 1.022 → 1.007 → 1.004 → 0.995 → 0.994.

The obvious reading is that small hospitals deliver worse heart-failure care. Part of it probably is that. But part of it is arithmetic, and the arithmetic matters more for anyone building a supervisor.

Why the smallest hospitals almost never reach the third column

Press dim everything that doesn’t survive and watch which ribbons stay lit. They come overwhelmingly from the bottom of the diagram. 27.8% of hospitals with 800 or more discharges clear the significance bar; only 12.1% of those under 100 do. A 90-discharge hospital’s confidence interval is wide enough to swallow most of the national spread.

There is a second, subtler thing happening. Look at the standard deviation of the published grades by band: 0.040 in the smallest, rising to 0.077 in the largest. Raw sampling noise would do the reverse — small samples scatter more, not less. The published numbers are tighter for small hospitals because CMS’s hierarchical model deliberately pulls uncertain estimates back toward the mean.

Which means the published grade is a compromise between what the hospital did and what the model believes about hospitals in general. One hospital in this file recorded 14 readmissions out of 31 heart-failure discharges — a 45% crude rate against a 19.4% expected rate, clearing p<.001 standing alone. Its published excess readmission ratio is 1.105. Shrunk nearly all the way back to normal.

One hospital in this file recorded 14 readmissions out of 31 heart-failure discharges — a 45% crude rate against a 19.4% expected rate, a result that clears p<.001 standing alone. Its published excess readmission ratio is 1.105. Shrunk nearly all the way back to normal.

80/20 — the builder read

Shrinkage is the right call for paying hospitals and the wrong call for catching a drifting agent. A supervisory system that pulls low-volume estimates toward the mean is, by construction, slowest to flag the single clinic where something has gone wrong — which is where agent failures will actually show up first. If you are building in this space, the product is not “a dashboard of ratios.” It is an explicit answer to: how many patients does this agent have to hurt before your monitor says so, and who gets paged at that threshold? Write that number down. Nobody selling agent-monitoring today will.

That is the shape of the problem ADVOCATE just funded someone to solve. Duke’s validation platform spans five health systems and rural sites; Kaiser’s deployment covers 21 medical centers and 260+ clinics. Stanford’s supervisor has to say something useful about all of them — including the ones that, on fourteen years of precedent, it will not be able to see.

Where this is thin