Medicare’s office-visit levels are the most heavily specified clinical judgment in America — written down, revised twice, audited, and worth real money to get right. Here is what half a million clinicians did with that spec.
Today’s newsletter is about a filmmaker who had taken hundreds of flights and could not say what flying was like until somebody asked. The claim underneath it: clinical judgment compresses into something faster than language, and the thing nobody can do now is write the spec — say in plain words what the decision is and when it should refuse to fire.
The obvious objection is that we already write specs all the time. So take the best case. Take the one clinical judgment America has tried hardest to specify.
“Was that office visit straightforward, low, moderate, or high complexity?” It has a written definition. It was rewritten from scratch in 2021 to be clearer. It has a decision table, worked examples, CME modules, coding consultants, and an audit apparatus. And it is attached to money, which is the strongest incentive to converge that a health system has.
If a decision this specified still can’t be pinned down, the bottleneck was never the writing.
Each dot below is one clinician who billed established-patient office visits to Medicare in 2024. Horizontal position is how many they billed. Vertical position is the share they coded at level 4 or 5 (99214/99215) rather than level 2 or 3 (99212/99213) — a single number standing in for “how complex do I think my visits are.”
One honesty note about that field: it is 90 clinicians drawn at random from each of sixteen specialties, which deliberately over-weights the small ones so the scatter is readable. It is a picture of the disagreement, not a national average — every number in the tables below is computed on the full population instead.
Push the min visits slider and the first thing that happens is the top and bottom rails empty out. That is the built-in trap in this dataset, and it is worth naming precisely rather than quietly filtering away.
CMS suppresses any provider–code row covering ten or fewer beneficiaries. A cardiologist with 400 level 4s and eight level 3s does not appear as 98% — the level 3 row is deleted, and she appears as 100%. Among clinicians with 11–49 visits, 79.9% have only one of the four codes present at all, which is why 85% of them sit at exactly 0% or exactly 100%. That bimodality is an artifact of the release, not of practice.
| Visit volume | clinicians | only 1 of 4 codes present | at 0% or 100% | median | middle half spans |
|---|---|---|---|---|---|
| 11–49 | 97,812 | 79.9% | 85.1% | 48% | 0% – 100% |
| 50–199 | 186,844 | 25.1% | 37.7% | 60% | 26% – 89% |
| 200–499 | 149,950 | 9.7% | 19.1% | 67% | 34% – 89% |
| 500–1,499 | 94,808 | 5.6% | 13.1% | 70% | 36% – 91% |
| 1,500+ | 11,558 | 4.0% | 9.8% | 65% | 27% – 90% |
So the spurious signal dissolves exactly the way a small-n artifact should: the rails collapse from 85% of clinicians to 10%, and the “only one code present” column falls from 80% to 4%.
And the spread underneath it does not move at all.
Among the 11,558 highest-volume clinicians in the country — people billing more than 1,500 established office visits a year, whose numbers cannot be a rounding artifact — the middle half still runs from 27% to 90% level 4/5. Two clinicians in that group, both busy, both audited, both reading the same guideline, can differ by sixty-three percentage points on the question was that visit complicated.
Below: every specialty with at least 1,000 clinicians billing 200+ established visits in 2024. The bar runs from the 10th to the 90th percentile of clinicians in that specialty; the notch is the median. This is within-specialty spread — same training, same patients, same code book.
It is not evidence of fraud, and the honest reading has to say so. Real drivers of a high level 4/5 share include an older and sicker panel, a referral-heavy practice, longer visit slots, and a scribe or coding team that documents more thoroughly. Two of the tightest, highest-median specialties here — endocrinology and interventional cardiology — are exactly the ones where nearly every visit genuinely is moderate complexity.
It is also not a clean measure of judgment. It is a measure of judgment plus documentation habit plus billing infrastructure, and those cannot be separated in claims data. A clinician who thinks in level 4s and writes in level 3s is invisible here.
But that is the point rather than a hole in it. Even the confounders are unwritten. Nobody can hand you the rule that says how much a scribe should shift your coding, or how sick a panel has to be before 90% moderate-complexity is right. Those are also reflexes. They are also faster than language.
The spec is the deliverable, and it is the part you can’t buy. America wrote out one clinical judgment in exhaustive detail, revised it, attached money to it, and got a 55-point interquartile range among the 149,950 clinicians billing 200–499 of these visits a year. Not because the writers were lazy — because a written rule can only ever be a lossy compression of a reflex.
Which is the whole reason the “learn to code” advice misses. Any model can write the Python. What no model can do is tell you why you went back and put your hand on that belly a second time. Take one case this week where you were right and don’t know why, and write down the reasoning you didn’t have time to have. That is the artifact. Everything else is downstream of it.