clinicians.build · interactive · 14 sep 2026

The Scoreboard the Agent Is Aimed At

ARPA-H has put $62.7 million into autonomous heart-failure agents, with a supervisory AI to watch them, and six health system leaders who cannot say who is liable when it goes wrong. The outcome those agents exist to move is the 30-day readmission rate. Here is that measure for every hospital Medicare grades on it — and how much of the grade survives a look at its own sample size.

Primary source: Becker’s Hospital Review, “Who will carry the liability? 6 health system leaders on autonomous AI’s biggest question,” 11 September 2026
Data: CMS Hospital Readmissions Reduction Program, FY 2026 file (discharges 1 Jul 2021 – 30 Jun 2024), via MIMI Labs

Six ADVOCATE awardees split the work. Atman Health, Tempus AI and Updoc build patient-facing heart-failure agents; Stanford builds a supervisory AI to watch them; Duke and Kaiser deploy. FDA package due in twenty-four months. Asked the obvious question, none of the six leaders Becker’s polled had an answer. The Christ Hospital’s Joy Oh put it best:

“If a supervisory agent fails to detect an inappropriate recommendation that results in patient harm, who is ultimately responsible: the developer of the worker agent or the developer of the supervisory agent?”

Before that question can be answered, a quieter one has to be: can you even tell? Liability runs on attribution, and attribution runs on measurement. So look at the measurement that already exists — the one CMS has been using to dock hospital payments since 2012, the one these agents are built to improve.

Across 2,316 hospitals, Medicare counted 858,817 heart-failure discharges and 169,065 readmissions inside thirty days over three years — a national rate of 19.7%. CMS grades each hospital with an excess readmission ratio: its predicted rate over the rate expected for its case mix. Above 1.00 is worse than expected. 1,187 hospitals are above it. 1,129 are below.

Every graded hospital, plotted against the size of its own sample

Each dot is one hospital. Horizontal is how many cases it had. Vertical is observed readmissions over expected — 1.00 means exactly what its case mix predicts. The shaded funnel is the range you would get from chance alone at that volume. Dots inside it are hospitals whose grade you cannot distinguish from a coin landing badly.

Condition graded — all six HRRP cohorts
all hospitals
Confidence required before a hospital counts as distinguishable
Hospitals shown
 
Worse beyond chance
 
Better beyond chance
 
Indistinguishable
 
Observed ÷ expected 30-day readmissions, by case volume
The same hospitals as a ranked strip — CMS excess readmission ratio, one tick each
worse beyond chance better beyond chance indistinguishable from its own expected rate · hover any dot · case axis is logarithmic

Three things the shape tells you

1. Most of the scoreboard is not a measurement. CMS sorts these hospitals almost exactly in half — 1,187 worse than expected, 1,129 better. Test each one against its own expected count and 338 of 2,316 (14.6%) differ from it beyond chance at p < 0.05. The other 1,978 — 85.4% — are indistinguishable. Tighten the bar to the three-sigma limit a funnel plot conventionally uses and only 60 hospitals (2.6%) survive.

2. The reason is volume, and the volumes are small. The median graded hospital had 274 heart-failure discharges in three years — about 91 a year. 849 hospitals (36.7%) had fewer than 200. Raise the minimum-cases slider and watch the grey swallow the edges: among hospitals with fewer than 200 cases, 10.4% are distinguishable; among those with 500 or more, 22.7%. Same measure, different amount of evidence.

3. CMS already knows this, which is the interesting part. Press show CMS’s ratio instead. The cloud collapses toward 1.00, because the excess readmission ratio is not a raw rate — it comes out of a hierarchical model that shrinks each hospital toward the national average in proportion to how little data it has. The raw ratios plotted here run from 0.45 to 2.32; CMS’s published ratios for the same hospitals run from 0.71 to 1.36. The statistical fix is already inside the measure. It is just invisible to everyone who reads the ranking.

Where this analysis is weak — and it is, in four specific places

If you build clinical AI that promises an outcome

The liability question in today’s briefing has a measurement question sitting underneath it, and the measurement question is answerable now.