America took one clinical judgment — how complicated was that visit? — and wrote it down. Definitions, a decision table, worked examples, audits, money attached. Press play and watch what the writing bought.
Today’s newsletter is about a man who had taken hundreds of flights and could not say what flying was like until an Amish woman who had never left Holmes County asked him. The argument underneath: expertise compresses into something faster than language, and the thing nobody can do now is write the spec.
The obvious objection is that we write specs all the time. So take the best case in American medicine. Take the judgment we tried hardest to pin down.
Each dot below is one clinician who billed established-patient office visits to Medicare in 2024. Left to right is the share they coded at level 4 or 5 — moderate or high complexity — rather than level 2 or 3.
They start where a working spec would put them: together. Press play.
Same rule. Same code book. Same audits. Sixty-seven points of daylight between the clinician at the 25th percentile and the one at the 75th.
One honesty note about that field: it is 90 clinicians drawn at random from each of sixteen specialties, which deliberately over-weights the small ones. It is a picture of the disagreement, not a national average. Every number in the tables below is computed on the full population instead.
Before drawing any conclusion from a spread, you should try to break it. This one has an obvious way to break: CMS deletes any provider–code row covering ten or fewer beneficiaries. A cardiologist with 400 level 4s and eight level 3s doesn’t appear at 98% — the level 3 row is gone and she appears at 100%.
So the low-volume clinicians pile up on the two walls. Step through the volume bands and watch the walls empty.
The national figures behind that, computed on all 540,972 clinicians rather than the sample:
| Visit volume | clinicians | only 1 of 4 codes present | at 0% or 100% | median | middle half spans |
|---|---|---|---|---|---|
| 11–49 | 97,812 | 79.9% | 85.1% | 48% | 0% – 100% |
| 50–199 | 186,844 | 25.1% | 37.7% | 60% | 26% – 89% |
| 200–499 | 149,950 | 9.7% | 19.1% | 67% | 34% – 89% |
| 500–1,499 | 94,808 | 5.6% | 13.1% | 70% | 36% – 91% |
| 1,500+ | 11,558 | 4.0% | 9.8% | 65% | 27% – 90% |
The artifact behaves exactly as an artifact should: the walls fall from 85% of clinicians to 10%, and the “only one code present” column collapses from 80% to 4%.
And the middle does not move. Among the 11,558 busiest clinicians in the country — more than 1,500 established office visits each, numbers far too large to be a rounding artifact, practices far too visible to skip an audit — the middle half still runs from 27% to 90%.
It isn’t fraud, and it would be cheap to imply otherwise. An older panel, a referral-heavy practice, longer slots, a scribe who documents more completely — all of those legitimately move this number. It also isn’t a clean measure of judgment: it’s judgment plus documentation habit plus billing infrastructure, and claims data cannot separate them.
But notice what just happened. Every one of those confounders is also unwritten. Nobody can hand you the rule for how much a scribe should shift your coding, or how sick a panel has to be before 90% moderate complexity is the right answer. Those are reflexes too.
Specification is lossy, and the loss is where the expertise lives. This is the most-specified judgment we have — written, revised, audited, paid for — and nationally it still comes out 55 points wide among the 149,950 clinicians billing 200–499 of these visits a year. Not because the drafters were careless. Because a written rule is a compression of a reflex, and compression drops things.
Which is why “learn to code” was always the wrong assignment. A model will write the Python. What no model can produce is the sentence explaining why you went back and put your hand on that belly a second time. Take one case this week where you were right and can’t say why, and write down the reasoning you didn’t have time to have. That is the deliverable. Everything else is downstream.