ARPA-H has put $62.7 million into autonomous heart-failure agents, with a supervisory AI to watch them, and six health system leaders who cannot say who is liable when it goes wrong. The outcome those agents exist to move is the 30-day readmission rate. Here is that measure for every hospital Medicare grades on it — and how much of the grade survives a look at its own sample size.
Six ADVOCATE awardees split the work. Atman Health, Tempus AI and Updoc build patient-facing heart-failure agents; Stanford builds a supervisory AI to watch them; Duke and Kaiser deploy. FDA package due in twenty-four months. Asked the obvious question, none of the six leaders Becker’s polled had an answer. The Christ Hospital’s Joy Oh put it best:
“If a supervisory agent fails to detect an inappropriate recommendation that results in patient harm, who is ultimately responsible: the developer of the worker agent or the developer of the supervisory agent?”
Before that question can be answered, a quieter one has to be: can you even tell? Liability runs on attribution, and attribution runs on measurement. So look at the measurement that already exists — the one CMS has been using to dock hospital payments since 2012, the one these agents are built to improve.
Across 2,316 hospitals, Medicare counted 858,817 heart-failure discharges and 169,065 readmissions inside thirty days over three years — a national rate of 19.7%. CMS grades each hospital with an excess readmission ratio: its predicted rate over the rate expected for its case mix. Above 1.00 is worse than expected. 1,187 hospitals are above it. 1,129 are below.
Each dot is one hospital. Horizontal is how many cases it had. Vertical is observed readmissions over expected — 1.00 means exactly what its case mix predicts. The shaded funnel is the range you would get from chance alone at that volume. Dots inside it are hospitals whose grade you cannot distinguish from a coin landing badly.
1. Most of the scoreboard is not a measurement. CMS sorts these hospitals almost exactly in half — 1,187 worse than expected, 1,129 better. Test each one against its own expected count and 338 of 2,316 (14.6%) differ from it beyond chance at p < 0.05. The other 1,978 — 85.4% — are indistinguishable. Tighten the bar to the three-sigma limit a funnel plot conventionally uses and only 60 hospitals (2.6%) survive.
2. The reason is volume, and the volumes are small. The median graded hospital had 274 heart-failure discharges in three years — about 91 a year. 849 hospitals (36.7%) had fewer than 200. Raise the minimum-cases slider and watch the grey swallow the edges: among hospitals with fewer than 200 cases, 10.4% are distinguishable; among those with 500 or more, 22.7%. Same measure, different amount of evidence.
3. CMS already knows this, which is the interesting part. Press show CMS’s ratio instead. The cloud collapses toward 1.00, because the excess readmission ratio is not a raw rate — it comes out of a hierarchical model that shrinks each hospital toward the national average in proportion to how little data it has. The raw ratios plotted here run from 0.45 to 2.32; CMS’s published ratios for the same hospitals run from 0.71 to 1.36. The statistical fix is already inside the measure. It is just invisible to everyone who reads the ranking.
Denver Health’s Daniel Kortsch, MD, told Becker’s the evidence under the AI-supervises-AI model covers four conditions, none of them heart failure, with a headline trial of 32 patients. The dashed navy line on the chart sits at 32 cases. At the national heart-failure rate, a 32-patient sample’s 95% interval spans 6.3% to 31.3% — an observed-over-expected range of 0.32 to 1.59. A trial that size cannot tell a hospital that cuts readmissions by a third from one that raises them by half. The smallest hospital CMS grades here has 31 cases, and CMS declines to publish a ratio below that. Whatever the ADVOCATE agents do to readmissions, the instrument that would catch it needs hundreds of patients per site, per year, for years.
The liability question in today’s briefing has a measurement question sitting underneath it, and the measurement question is answerable now.