clinicians.build · interactive · august 9, 2026

The Level Four Line

Medicare’s office-visit levels are the most heavily specified clinical judgment in America — written down, revised twice, audited, and worth real money to get right. Here is what half a million clinicians did with that spec.

Primary source: Barbara Hays & Cindy Hughes, Coding Level 4 Office Visits Using the New E/M Guidelines
AAFP Family Practice Management · January/February 2021
Grounding data: MIMI Labs · CMS Medicare Physician & Other Practitioners PUF, service year 2024

Today’s newsletter is about a filmmaker who had taken hundreds of flights and could not say what flying was like until somebody asked. The claim underneath it: clinical judgment compresses into something faster than language, and the thing nobody can do now is write the spec — say in plain words what the decision is and when it should refuse to fire.

The obvious objection is that we already write specs all the time. So take the best case. Take the one clinical judgment America has tried hardest to specify.

“Was that office visit straightforward, low, moderate, or high complexity?” It has a written definition. It was rewritten from scratch in 2021 to be clearer. It has a decision table, worked examples, CME modules, coding consultants, and an audit apparatus. And it is attached to money, which is the strongest incentive to converge that a health system has.

If a decision this specified still can’t be pinned down, the bottleneck was never the writing.

Every clinician, one dot, one number

Each dot below is one clinician who billed established-patient office visits to Medicare in 2024. Horizontal position is how many they billed. Vertical position is the share they coded at level 4 or 5 (99214/99215) rather than level 2 or 3 (99212/99213) — a single number standing in for “how complex do I think my visits are.”

1,440 clinicians · random sample, 90 per specialty · CMS Physician & Other Practitioners PUF, 2024
≥ 11
clinicians shown
median lvl 4/5 share
middle-half spread
pinned at 0% or 100%

One honesty note about that field: it is 90 clinicians drawn at random from each of sixteen specialties, which deliberately over-weights the small ones so the scatter is readable. It is a picture of the disagreement, not a national average — every number in the tables below is computed on the full population instead.

The part of the spread that isn’t real

Push the min visits slider and the first thing that happens is the top and bottom rails empty out. That is the built-in trap in this dataset, and it is worth naming precisely rather than quietly filtering away.

CMS suppresses any provider–code row covering ten or fewer beneficiaries. A cardiologist with 400 level 4s and eight level 3s does not appear as 98% — the level 3 row is deleted, and she appears as 100%. Among clinicians with 11–49 visits, 79.9% have only one of the four codes present at all, which is why 85% of them sit at exactly 0% or exactly 100%. That bimodality is an artifact of the release, not of practice.

So the spurious signal dissolves exactly the way a small-n artifact should: the rails collapse from 85% of clinicians to 10%, and the “only one code present” column falls from 80% to 4%.

And the spread underneath it does not move at all.

Among the 11,558 highest-volume clinicians in the country — people billing more than 1,500 established office visits a year, whose numbers cannot be a rounding artifact — the middle half still runs from 27% to 90% level 4/5. Two clinicians in that group, both busy, both audited, both reading the same guideline, can differ by sixty-three percentage points on the question was that visit complicated.

Where each specialty sits

Below: every specialty with at least 1,000 clinicians billing 200+ established visits in 2024. The bar runs from the 10th to the 90th percentile of clinicians in that specialty; the notch is the median. This is within-specialty spread — same training, same patients, same code book.

31 specialties · 10th–90th percentile of clinician-level 4/5 share · clinicians with ≥200 visits
Podiatry’s median is 6.6% and endocrinology’s is 95.1%. That gap is real and defensible — a diabetic foot check and a poorly controlled type 1 visit are not the same encounter. The gap the spec was supposed to close is the horizontal length of each bar — and the shortest bar on the board, endocrinology’s, is still thirty points wide.

What this is and isn’t evidence of

It is not evidence of fraud, and the honest reading has to say so. Real drivers of a high level 4/5 share include an older and sicker panel, a referral-heavy practice, longer visit slots, and a scribe or coding team that documents more thoroughly. Two of the tightest, highest-median specialties here — endocrinology and interventional cardiology — are exactly the ones where nearly every visit genuinely is moderate complexity.

It is also not a clean measure of judgment. It is a measure of judgment plus documentation habit plus billing infrastructure, and those cannot be separated in claims data. A clinician who thinks in level 4s and writes in level 3s is invisible here.

But that is the point rather than a hole in it. Even the confounders are unwritten. Nobody can hand you the rule that says how much a scribe should shift your coding, or how sick a panel has to be before 90% moderate-complexity is right. Those are also reflexes. They are also faster than language.