The AI-versus-physician argument assumes the comparator is uniform. It isn’t — and the public data that is supposed to show you the spread cannot actually resolve it. Here is every US hospital’s hip and knee complication rate, with the error bars turned on.
AMA chief executive John Whyte and Zeke Emanuel spent this weekend arguing on video about whether AI can outperform physicians, off the back of an August JAMA analysis claiming it can on many clinical tasks. Buried further down the same piece was Garner Health’s Nick Reber, pointing out a roughly four-fold spread in complication rates between best- and worst-quartile physicians inside a single brand-name health system.
That is the more interesting number, because it breaks the frame. “AI versus the physician” assumes there is a physician — one comparator, with one performance level. If the variance inside the profession is larger than the variance being argued about, the whole debate is being conducted about the wrong quantity.
So let’s try to see the spread. CMS publishes, for every hospital that does enough of them, the risk-standardised rate of complications after elective hip and knee replacement — one of the cleanest, most-studied, most-standardised procedures in American medicine. If the four-fold spread is visible anywhere in public data, it is here.
It is. And then it isn’t.
Each vertical mark is one hospital, sorted worst to best. Let it run.
Stage one looks like a finding. A smooth gradient from about 1.4% to about 9.3% — a six-fold range across American hospitals on a single, well-defined procedure. Print it, rank it, put it in a board deck.
Stage two turns on the 95% confidence interval CMS publishes alongside each rate. The typical interval is about 4.2 percentage points wide. The entire distribution, from the 5th to the 95th percentile, is about 2.3 points wide. The uncertainty on one hospital’s number is nearly twice the size of the thing the number is supposed to distinguish.
Stage three is CMS agreeing. Of the 1,655 hospitals with a publishable rate, CMS classifies 97.6% as “no different than the national rate.” Twenty-eight are better. Twelve are worse. The other 1,615 are a rank ordering of noise, published with four significant figures of confidence.
The median hospital in this file contributed about 75 eligible cases. At a 3.5% complication rate that is roughly two or three events. Nothing you build on top of a two-event denominator will survive contact with next year’s data — and a leaderboard is the most confident possible presentation of a two-event denominator.
Drop the small hospitals and watch the tails come off. Almost everything at both extremes of the ranking is a low-volume facility; the “worst hospital in America” and the “best hospital in America” are, overwhelmingly, the hospitals that did the fewest operations.
This is the same failure as the drug that moved the biomarker and changed nothing, and the same failure as the model that hits 92% top-3 accuracy on a benchmark nobody has connected to an order being placed.
A number that cannot resolve the thing it names is not a measurement. It is a ritual. Before you route, rank, benchmark or gate on any published quality metric, ask what its denominator is and how wide its interval is — and if the interval is wider than the spread, you are building a product on top of a coin flip with a decimal point.