“Clinical judgment requires deciding when low risk is low enough,” writes Graham Walker. That call is usually invisible. Screening mammography is the one place Medicare puts a public number on it: the share of women screened who get called back for more imaging. Here is every hospital that reports it (…), against the 5–12% window radiology’s own guideline calls acceptable.
Primary source:Graham Walker, MD, thesis statement, Every (“AI makes uncertainty cheap to explore without making it cheap to resolve”). Data: CMS Care Compare Outpatient Imaging Efficiency – Hospital (file updated July 22, 2026), measure OP-39 Breast Cancer Screening Recall Rates and OP-10 Abdomen CT Use of Contrast Material, Medicare outpatients, July 1, 2024 – June 30, 2025. Year-over-year persistence computed on the two prior reporting periods via MIMI Labs. Target window: OP-39 measure specification (CBE #4220).
Four states break the chart, and it’s probably the ruler
The national median is …. Connecticut’s median hospital calls back …, New York’s …. Judgment doesn’t cluster that neatly at state lines. OP-39 counts any breast ultrasound, MRI or diagnostic mammogram within 45 days of a screen, same day included. Connecticut was the first state to require dense-breast notification and coverage of supplemental ultrasound (2009); a routine same-visit ultrasound for dense tissue is counted by this measure as a “recall.” That is the likeliest explanation, not a proven one. Click Hide CT, NY, NJ, FL and the share above 12% falls from … to ….
It’s a habit, not a bad year
Outside those four states, a hospital’s recall rate in one reporting year predicts the next (correlation 0.71 across 2,891 hospitals, July 2022–June 2023 vs. July 2023–June 2024). CMS doesn’t publish how many screens each hospital read, so this can’t be a funnel plot. Year-to-year stability is the best available evidence that the spread isn’t just small-sample noise.
“Low enough” isn’t a hospital personality
Switch to Callbacks vs. double CT scans. OP-10 is the share of abdominal CTs done both with and without contrast, which is usually a second scan to settle a question that didn’t need settling. If some hospitals were simply “can’t-stop” places, the two would line up. They don’t: correlation … across … hospitals reporting both. The threshold lives in a department and a reading room, one clinician at a time. That matches Walker’s point: nerve is built one shift at a time.
The window has two edges
Below 5% isn’t automatically a virtue. … of hospitals sit there. A very low recall rate can mean a confident reader or cancers that nobody called back. This file has no cancer-detection data, so it can’t tell which. “Low enough” is a decision with a cost on both sides, and that’s why it needs a name on it.
Where this is thin — read before quoting
Claims, not reads. OP-39 is built from Medicare billing. Radiologists’ own registry numbers correlate only moderately with it (r = 0.43, Lee et al., AJR 2018), and the claims version ran lower on average.
No denominators, no outcomes. Care Compare publishes the rate only. There’s no volume, no cancer detection rate and no biopsy yield, so a dot can’t say whether its callbacks found anything.
Medicare hospital outpatients only. Freestanding imaging centers and commercially insured women under 65 aren’t here.
A dot outside the window is a question, not a verdict. Patient mix (prior cancers, implants, first-ever screens) moves recall legitimately.