Thirteen U.S. physicians read 750 fetal ultrasound stills twice — once with an FDA-cleared assistant, once with nothing. Everyone will quote the 22-point AUC lift. This is the other arm: what the same readers, on the same images, did alone.
750 images. Switch the AI off and 114 findings go dark.
Every square is one still image the readers saw. The left block is the 250 images that contained one of eight abnormal findings — absent cavum septum pellucidum, absent corpus callosum, thoracic situs inversus and five others. The right block is the 500 normal images. Red means a flag was raised. Solid red on the left is a catch; a hollow square is an abnormality that went by.
abnormality caught abnormality missed normal image falsely flagged normal image called normal
caught
—
missed
—
false flags on normals
—
mean AUC across 13 readers
—
Published aggregates, laid out as a field: sensitivity 54.2% → 88.5%, specificity 80.7% → 87.9%, over 250 abnormal and 500 normal images. Which individual square flips is illustrative — the paper reports pooled rates, not per-image results. The counts are exact.
80/20 — why the baseline is the finding
The lift is a product claim. The baseline is a fact about the world. "90.9% AUC" is unreadable on its own — 90.9% against what? The same paper's 68.9% unassisted arm is what makes the first number mean anything, and it is the arm that almost never survives into a datasheet. Before you evaluate any clinical AI tool this month, ask what the unassisted arm was and who ran it. No answer means no performance claim — just a demo with a percentage attached.
And look at the 26%. Unassisted, three trained specialties — maternal-fetal medicine, OB/GYN, radiology — agreed with each other on roughly a quarter of images. That is not a statement about these readers. It is a statement about how hard a single frozen frame is with no sweep, no history and no second view. We only know it because somebody ran the arm.
the stress test the study didn't run
A third of these images were abnormal. Your clinic isn't.
Enriched prevalence is standard for reader studies and it is also the thing that makes sensitivity look load-bearing. Drag prevalence down toward what an anomaly scan actually finds and watch what happens to the number a clinician actually experiences — the chance that a flag means anything.
PPV unassisted
—
PPV with AI
—
false flags per true catch, with AI
—
Bayes on the paper's own published sensitivity and specificity. At the study's 33.3% prevalence, AI moves PPV from 58.4% to 78.5%. At a 3% prevalence — closer to a real second-trimester anomaly screen — the same rates give 8.0% unassisted and 18.4% with AI. The tool more than doubles the value of a flag and still leaves roughly four false alarms for every real finding, because specificity is what governs at low prevalence and 87.9% of 500 normals is a lot of normals.
😤 "This is a vendor study." It is, completely. Five authors are full-time Sonio employees, four more sit on Sonio's scientific advisory board, and Sonio funded the work — there is no independent name on the byline. That doesn't make the numbers fake. It means the design choices — retrospective, still frames, 250 abnormal of 750 — were made by the party with a stake in the answer.
😤 "So a machine beat doctors, again." No. A machine beat doctors doing a task nobody actually does: one frozen frame, no sweep, no history, no repeat view. It is still interesting, because the thing being measured is exactly what the AI sees too — but a 40-second look at a still is not a scan.
😤 "The specificity went up, so the false-alarm worry is overblown." It did go up, 80.7% to 87.9%, and that is genuinely the harder result to get. The prevalence panel is not saying the tool is worse; it is saying the number that decides whether a clinician trusts a flag is not in the abstract, and it moves by 60 points depending on a parameter the study fixed by design.