clinicians.build · interactive · 8 sep 2026

Four Times Better, Give or Take

The AI-versus-physician argument assumes the comparator is uniform. It isn’t — and the public data that is supposed to show you the spread cannot actually resolve it. Here is every US hospital’s hip and knee complication rate, with the error bars turned on.

Primary source: Christina Farr & Annalise Merelli, “Doctors get into the ring about AI’s potential in healthcare,” Second Opinion, 7 Sep 2026
Data: CMS Care Compare, provider data release of 1 May 2026 · 1,655 hospitals with a reportable rate
queried via MIMI Labs · every statistic on this page recomputed in-browser from the values below

AMA chief executive John Whyte and Zeke Emanuel spent this weekend arguing on video about whether AI can outperform physicians, off the back of an August JAMA analysis claiming it can on many clinical tasks. Buried further down the same piece was Garner Health’s Nick Reber, pointing out a roughly four-fold spread in complication rates between best- and worst-quartile physicians inside a single brand-name health system.

That is the more interesting number, because it breaks the frame. “AI versus the physician” assumes there is a physician — one comparator, with one performance level. If the variance inside the profession is larger than the variance being argued about, the whole debate is being conducted about the wrong quantity.

So let’s try to see the spread. CMS publishes, for every hospital that does enough of them, the risk-standardised rate of complications after elective hip and knee replacement — one of the cleanest, most-studied, most-standardised procedures in American medicine. If the four-fold spread is visible anywhere in public data, it is here.

It is. And then it isn’t.

1,655 hospitals, ranked

Each vertical mark is one hospital, sorted worst to best. Let it run.

Stage 1 of 3 — the published rate
hover any hospital
CMS: worse than national CMS: better than national CMS: no different
Hospitals
reportable rate
Worst ÷ best
published rates
5th–95th spread
percentage points
Typical error bar
mean 95% CI width
“No different”
CMS’s own verdict

What just happened

Stage one looks like a finding. A smooth gradient from about 1.4% to about 9.3% — a six-fold range across American hospitals on a single, well-defined procedure. Print it, rank it, put it in a board deck.

Stage two turns on the 95% confidence interval CMS publishes alongside each rate. The typical interval is about 4.2 percentage points wide. The entire distribution, from the 5th to the 95th percentile, is about 2.3 points wide. The uncertainty on one hospital’s number is nearly twice the size of the thing the number is supposed to distinguish.

Stage three is CMS agreeing. Of the 1,655 hospitals with a publishable rate, CMS classifies 97.6% as “no different than the national rate.” Twenty-eight are better. Twelve are worse. The other 1,615 are a rank ordering of noise, published with four significant figures of confidence.

Push the volume slider

Drop the small hospitals and watch the tails come off. Almost everything at both extremes of the ranking is a low-volume facility; the “worst hospital in America” and the “best hospital in America” are, overwhelmingly, the hospitals that did the fewest operations.

all 1,655
Hospitals shown
after filter
Worst ÷ best
collapses as n rises
Mean error bar
points, 95% CI

Where this is thin

The builder version

This is the same failure as the drug that moved the biomarker and changed nothing, and the same failure as the model that hits 92% top-3 accuracy on a benchmark nobody has connected to an order being placed.

A number that cannot resolve the thing it names is not a measurement. It is a ritual. Before you route, rank, benchmark or gate on any published quality metric, ask what its denominator is and how wide its interval is — and if the interval is wider than the spread, you are building a product on top of a coin flip with a decimal point.