A network meta-analysis of 53 studies and 7 million admissions puts the best sepsis-prediction models at an AUROC of 0.88. Then the number nobody puts on a slide: pooled positive predictive value 34.2%. Here is what that means at your denominator — and what 3,106 US hospitals look like when you plot them against theirs.
Discrimination is not the same as a workable alert. An AUROC of 0.88 says the model ranks septic patients above non-septic ones most of the time. A positive predictive value of 34.2% says that when it fires, it is wrong about twice for every once it is right. The paper names the consequence itself: a substantial risk for alarm fatigue.
The gap between those two numbers is not a modelling problem. It is arithmetic, and the variable that drives it is the one your vendor's slide never contains: how common sepsis actually is in the population you are screening.
If you can't state false alerts per nurse per shift, you don't have a deployment plan. You have a benchmark.
The meta-analysis reports heterogeneity above 95% and a 95% prediction interval from −0.06 to 0.30 — a polite way of saying the pooled estimate may not describe your hospital at all. So here is every US hospital that reported CMS sepsis bundle compliance, plotted against the thing that decides how much to believe it.
Each dot is one hospital. Horizontal is the denominator — how many severe sepsis and septic shock cases the measure was computed on, from 11 to 2,313, on a log scale. Vertical is the score. Drag the denominator floor and watch the shape of the cloud change.
Set the floor to zero and the SEP-1 cloud has hospitals at 0% and hospitals at 100%. Drag the floor to 100 cases and every one of them is gone — not because those hospitals improved, but because all 11 of them had denominators under 100. Standard deviation falls from 22.2 points below 25 cases to 13.9 points at 250 and above. The tails of this dataset are not performance. They are arithmetic on small numbers.
Before you quote an AUROC to anyone, compute PPV at your own site's prevalence and threshold. And before you rank yourself against a peer, check whether the gap survives their denominator.