Article
The Checkpoint That Caught What the Model Missed
A sepsis-prediction model sits one design choice away from the exact failure documented elsewhere in this series: an estimate that quietly stops informing a decision and starts being one. Two health systems built the same category of tool and avoided that failure, not with a more sophisticated model
- Published
- August 13, 2026
- Updated
- August 19, 2026
- Reading time
- 11 min

The Global Signal
Duke Health's Sepsis Watch, developed by the Duke Institute for Health Innovation, is a predictive model that flags patients at elevated risk of sepsis before clinical symptoms are obvious, directly confirmed to predict sepsis a median of five hours before clinical presentation. Duke's own program documentation states the system "doubled" its three-hour sepsis-bundle compliance rate as reported quarterly to the Centers for Medicare and Medicaid Services; a peer-reviewed account in Annals of Emergency Medicine puts this specifically at roughly 28 percent (2016 through the third quarter of 2018) rising to roughly 63 percent (fourth quarter of 2018 through the first quarter of 2019), a figure this draft treats as corroborated by a named peer-reviewed source rather than independently re-confirmed against the original journal page, which returned an access error on this pass. A 2025 peer-reviewed external validation, published in npj Digital Medicine, tested the model across four additional health systems and reported strong discriminative performance, with area-under-curve scores ranging from 0.906 to 0.960.
A separate, independently developed system, Johns Hopkins' TREWS (Targeted Real-time Early Warning System), led by researcher Suchi Saria and published in Nature Medicine, was associated with patients being 20 percent less likely to die of sepsis across a deployment spanning five hospitals, more than 4,000 clinicians, and roughly 590,000 patients, identifying 9,805 sepsis cases at 82 percent sensitivity, compared with under 50 percent for the tools it replaced.
Reduction in sepsis mortality associated with Johns Hopkins' TREWS deployment
Reported across a deployment spanning five hospitals, more than 4,000 clinicians, and roughly 590,000 patients; published in Nature Medicine.
The decision clarity here is not a smarter algorithm. It is a mandatory human checkpoint that was decided in advance.
What changes when the override is mandatory, not optional
A compliance or accuracy statistic cannot tell you whether a model's flag is confirmed by a real person or clicked through as a formality.
Naming who must confirm a model's output, and within what window, built into the workflow itself, is what actually distinguishes a governed deployment from one that is one incident away from failure.
Why the Visible Metric Misleads
The compliance-rate improvement, from roughly 28 percent to roughly 63 percent at Duke, is the number most likely to be repeated in any summary of this program's success, and it is real and externally validated. It is not, on its own, proof that the decision-governance structure behind it is sound. A hospital could post a similar compliance improvement using a system with no mandatory human confirmation at all, simply because the model's predictions were accurate enough on average, and still be exactly as exposed as an unsupervised claims-denial algorithm the first time the model was wrong in a way nobody was positioned to catch. The fact that actually distinguishes Duke's and Johns Hopkins' systems from a decision-blindness failure is not in either program's topline compliance or mortality statistic; it is the specific, less publicized design requirement that a named clinician confirm the flag within a defined window before anything happens to the patient.
The Leadership Move
The right move, demonstrated in both cases, is not choosing a more accurate model over a less accurate one. It is deciding, before deployment, exactly who is required to confirm a model's flag, within what time window, and building that requirement into the clinical workflow itself rather than leaving it as an optional step a busy clinician can skip.
- Ownership
Duke named a specific clinical lead, Dr. Cara O'Brien, alongside the Duke Institute for Health Innovation's data science leadership, and built a mandatory nurse-confirmation step directly into the model's deployment. Johns Hopkins built a defined acknowledgment window into TREWS itself. In both cases, ownership of the override decision was assigned to a specific clinical role before the system went live, not left to be worked out informally once it was in use.
- Tradeoff
Requiring a nurse's confirmation, or a clinician's acknowledgment within a set window, is slower than letting a validated model act on its own prediction, and that speed cost is real, especially in a condition where minutes matter. Both systems accepted that tradeoff deliberately, choosing a slightly slower, human-confirmed activation over a faster, unsupervised one, and the reported outcomes, a five-hour median lead time and a 20 percent reduction in sepsis mortality, show that tradeoff did not come at the expense of the underlying goal.
- Human consequence
Patients at both health systems benefited from earlier detection precisely because a human being remained the last check before a treatment protocol activated, not despite it. The mandatory confirmation step is also where a clinician's judgment about a specific patient's circumstances, something a population-trained model cannot see, gets a chance to matter before the model's prediction becomes an action.
Implication for Operators
Any organization deploying a predictive model for a consequential clinical or operational decision should treat the mandatory human-confirmation step, who is required to confirm, within what window, built into the workflow rather than left optional, as the actual measure of whether the system is governed soundly, not the model's own accuracy statistic or a topline compliance number. A model can be highly accurate and still be one unsupervised deployment decision away from becoming exactly the kind of decision-blindness failure this series documents elsewhere.
The same industry, healthcare, and the same category of tool, a predictive model applied to a consequential clinical decision, produced both a documented decision-blindness failure elsewhere in this series and the decision-clarity success this article describes. The difference was not model sophistication. It was a specific, unglamorous structural choice: a named clinical role, a required confirmation, a defined window, decided before the system ever touched a patient.
What made the difference was not a smarter algorithm. It was a mandatory human checkpoint decided in advance, rather than left to erode once the system was live.
Before automating a workflow or routing rule, how explicitly does your organization decide what the system should optimize for and who owns exceptions?
FAQ
Does requiring a nurse's confirmation slow down sepsis treatment?
It adds a deliberate step compared with fully automated activation, but the reported outcomes, a five-hour median lead time on detection at Duke and a 20 percent reduction in sepsis mortality associated with Johns Hopkins' TREWS, show that the confirmation step did not prevent faster, earlier intervention than prior approaches achieved.
Is this decision-support or autonomous diagnosis?
Decision-support. Neither system is designed or reported to make an unreviewed clinical determination; both require a specific clinical role to confirm the model's flag before a treatment protocol proceeds.
What would it look like if a hospital deployed a similar model without this safeguard?
The confirmation step would likely become a formality over time, clicked through rather than genuinely reviewed, the same pattern that allowed an unsupervised predictive model elsewhere in this series to function as an unreviewed coverage decision rather than an input to one.
How externally verified are these results?
Duke's Sepsis Watch has been externally validated across four additional health systems in a 2025 peer-reviewed study published in npj Digital Medicine . Johns Hopkins' TREWS outcomes were published in Nature Medicine , a peer-reviewed journal, based on a deployment across five hospitals.
What is the single most replicable element of this approach for another organization?
Naming, before deployment, exactly who is required to confirm a model's output and within what time window, and building that requirement into the workflow itself, rather than assuming a model's accuracy alone is sufficient governance.
Related intelligence
Article
The Data One Company Published and the Other Didn't
Two autonomous-vehicle companies experienced the same category of incident: a real-world edge case their systems had not been explicitly designed to handle. One tried to manage the disclosure. The other built continuous, third-party-audited disclosure into how it operates, and that second choice is
Article
The Algorithm and the Randomized Trial
Two education systems built algorithms to help make decisions about individual students at scale. One imposed its algorithm nationally, on results day, with no pilot and no override. The other tested its tool for years, on a smaller population, before trusting it broadly, and built it to route stude
Article
The Report That Audited Robots and Hallucinated Its Own Citations
A firm whose entire service is verification can still fail to verify its own output, if nobody explicitly owns checking AI-generated content before a client sees it. The irony sharpens considerably when the report in question is reviewing a government's own automated penalty system.