AI Risk & Governance
Low Risk Is Not No Risk: Using AI Safely in ED Disposition

Why this matters
AI risk scores can help emergency physicians see patterns sooner, but low risk is not a disposition plan.
Recommended next step
Pair this article with the free guide or course store if you want a more structured framework you can apply at the bedside or in leadership conversations.
At 3:00 a.m., a risk model labels a patient “low risk.” The number is clean. The dashboard is confident. The patient is still sitting in front of you with a story that does not fit the number.
That is the moment when clinical judgment matters most.
I am not opposed to AI in the emergency department. I am opposed to pretending that a prediction is a disposition plan. Admission, observation, transfer, and discharge are decisions about a real person in a real environment, with missing data, changing physiology, family concerns, access barriers, and consequences the model cannot carry.
A good risk tool can help me see a pattern sooner. It cannot make the final decision for me. Here is the practical framework I would use before allowing an AI score to influence disposition.
A risk score answers one question, not the whole disposition question
Every model has a target. It may estimate 31-day mortality, likelihood of admission, deterioration, or another outcome. That target is not identical to “safe to go home.”
The recent MARS-ED randomized trial is a useful warning. In 1,303 emergency-department patients, the RISK INDEX predicted 31-day mortality with an AUROC of 0.84, compared with 0.73–0.76 for clinical intuition. Yet treatment plans changed in only 1 of 644 patients who had access to the score—0.16%—and the investigators reported no change in clinical outcomes. The conclusion was not that prediction is useless. It was that prognostic accuracy alone does not guarantee clinical impact.
That is the first rule: name the outcome the model predicts, then separately name the decision you want the clinician to make. If those are different, the gap is where judgment, context, and safety checks belong.
“Low risk” is a probability statement, not a patient description
Risk scores compress a patient into a probability. Emergency physicians decompress the patient back into a person.
The score may not know that the patient lives two hours from the nearest hospital, cannot obtain the prescribed medication, has no one at home overnight, or is returning because the pain is worse despite following the plan. Those details are not noise. They change what a safe disposition means.
This is why I keep coming back to a point made in my Medium article on what AI still cannot do in emergency medicine: the texture of the story often carries the signal. Current tools can be excellent at pattern recognition and still miss the family member who describes a sudden change, the patient who has stopped behaving like themselves, or the second diagnosis that is less likely but dangerous.
The FDA’s clinical decision support guidance makes the same principle concrete. Software intended to support clinicians should enable them to independently review the basis of a recommendation rather than rely primarily on it. That means showing intended use, relevant inputs, validation population, known limitations, and patient-specific information—including missing or unexpected data.
The bedside override should be designed before go-live
A human override is not a failure of implementation. It is part of the implementation.
Before a model reaches clinicians, write the override pathway in plain language. Answer five questions:
- What decision does the score inform? Not “risk stratification.” Say “whether to consider observation,” “whether to repeat a reassessment,” or “whether to ask the attending to review the discharge plan.”
- Who must review it? Name the role, not an abstract “care team.” Is it the treating physician, charge nurse, triage clinician, or observation attending?
- What information must be checked? Include missing vitals, changing symptoms, medication access, baseline function, social support, and the patient’s own understanding of the plan.
- What findings require an override? A new symptom, discordant exam, unreliable follow-up, unexpected subgroup, or clinician concern should be a legitimate reason to step outside the score.
- Who owns the review when the first person is busy? A safety rule that depends on catching the right person at the right moment is not a safety rule.
The NIST AI Risk Management Framework calls for documented human-oversight roles, ongoing monitoring, and mechanisms to supersede or deactivate systems that perform outside intended use. In the ED, that translates into an obvious override button, a short reason list, a named operational owner, and a process for reviewing patterns of overrides.
Measure whether the tool changes care safely
A vendor will usually lead with discrimination metrics. Ask for them, but do not stop there.
At minimum, the local team should know sensitivity, specificity, positive predictive value, and negative predictive value in the population where the tool will run. Performance should be examined by age, race, sex, language, insurance status, arrival mode, acuity, and site when those factors affect the workflow. “Validated on millions of patients” is not a local validation plan.
The evidence base is a reminder to stay humble. A systematic review included 23 studies representing more than 16 million patients; it found promising discrimination, but most studies had high risk of bias because of limited external validation, low event rates, or inadequate calibration reporting. A model can look impressive on an area-under-the-curve chart and still be poorly calibrated for the patient in front of you.
Then measure what happens after implementation. Track:
- How often the score is available when needed.
- How often clinicians accept, ignore, or override it.
- Which patient groups generate disproportionate overrides.
- Whether the tool changes observation, admission, transfer, or discharge decisions.
- Whether return visits, diagnostic delay, and patient harm move in the right direction.
- Whether the score creates extra clicks, alert fatigue, or false reassurance.
If the model never changes a decision, it may be an expensive dashboard. If it changes every decision, that is not automatically success; it may be automation bias. The goal is not maximum adoption. The goal is better decisions with a visible safety net.
Disposition is where governance becomes clinical
The 2026 consensus statement from leading emergency medicine organizations affirms that emergency physicians retain authority for patient-care decisions and that AI should enhance, not replace, clinical judgment. That is more than a policy sentence. It is an operating requirement.
Disposition carries consequences outside the model’s target. Discharge includes follow-up reliability, transportation, medication access, caregiver support, and the patient’s ability to return if the condition changes. Admission consumes scarce capacity but may be safer when uncertainty is high; observation may be right because the model cannot resolve the uncertainty yet.
The physician also has to recognize distribution shift. A tool trained in one health system may be less reliable in a rural ED, a safety-net population, an older cohort, or a department with different documentation habits. If subgroup performance is unknown, that is not a footnote. It is a reason to lower the level of automation.
My practical threshold is simple: the closer the tool gets to a disposition decision, the more transparent, locally validated, interruptible, and physician-owned it must be. A score that helps prioritize chart review can tolerate more uncertainty than a score that quietly steers a discharge queue.
Dr. Chet’s Take
Use AI to widen your field of view, not to narrow your responsibility.
When a model says “low risk,” I want to know what it measured, what it missed, and whether the patient in front of me resembles the patients on which it was validated. I want the score next to the relevant data, not floating above it as an answer. I want the system to make it easy to disagree and hard to pretend that disagreement is a workflow failure.
The safest ED is not the one with the most AI. It is the one where clinicians know what the tool can do, what it cannot do, and who is accountable when the patient does not behave like the training set.
Key Takeaways
- A model’s target—mortality, admission, or deterioration—is not the same thing as a safe disposition.
- Treat “low risk” as a probability to investigate, not a discharge order.
- Design the human override, ownership, and escalation pathway before go-live.
- Validate locally and monitor calibration, subgroup performance, overrides, alert burden, and patient outcomes.
- Keep emergency physicians accountable for disposition decisions; AI should inform judgment, not replace it.
FAQ
Can an AI risk score decide whether I discharge a patient?
No. A risk score can support review, but disposition also depends on clinical trajectory, examination, follow-up reliability, patient preferences, and information the model may not capture.
What should I ask a vendor before piloting a tool?
Ask for intended use, training and validation populations, reference standard, sensitivity, specificity, predictive values, subgroup performance, missing-data behavior, and post-go-live monitoring. Ask what happens when the model is wrong.
Is a clinician override a sign that the model failed?
Not necessarily. An override may be the correct use of a tool when bedside findings, patient context, or an out-of-distribution case adds information the model does not have. Review override patterns for learning, not punishment.
What is the first safe use case for AI in disposition workflows?
Start with a low-autonomy task such as surfacing missing information, organizing a chart for physician review, or identifying patients who need a second look. Do not begin by allowing a score to automatically route patients toward discharge.
How do we know if the tool is helping?
Predefine success measures before the pilot: decision quality, safety outcomes, calibration, subgroup performance, overrides, alert burden, workflow time, and patient-centered outcomes. Accuracy alone is not enough.
If you're an emergency physician (or any clinician treating patients daily) trying to understand how AI will actually impact your clinical practice -- not just the hype -- I put together a free practical guide. You can download it here. AI in EM Survival Guide
Read more from Dr. Shermer on Medium
Sources
- U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- van Dam PMEL, et al. Machine learning for risk stratification in the emergency department (MARS-ED): a randomized controlled trial
- Kareemi H, et al. Machine Learning Versus Usual Care for Diagnostic and Prognostic Prediction in the Emergency Department: A Systematic Review
- American College of Emergency Physicians. Leading EM Organizations Issue Consensus Statement on Artificial Intelligence in EM
Keep reading
Related reading and your next step.
Ready to go further? Move from this article into structured training, scenario-based rehearsal, and more physician-written guidance.
Course
Translate the article into a repeatable framework
Use the physician-led course when you want a structured framework for evaluating AI tools, protecting clinical judgment, and leading implementation decisions.
Simulation
Practice the decision path under pressure
Use EM-Sim when you want scenario-based repetition that turns article-level insight into physician-facing emergency-medicine reps.
Blog
Browse more articles
Explore the full blog for more on AI in emergency medicine, then head to the course and simulation pages when you want the structured next step.
Related Articles
AI in Emergency Medicine
Before Your ED Buys AI, Demand the Evidence Packet
A vendor demo is not evidence that an AI tool will work in your emergency department. Here is the evidence packet and local stop rule to require before go-live.
AI Risk & Governance
AI Change Control in the ED: Stop Model Drift Before It Reaches the Bedside
A practical ED playbook for tracking AI changes, detecting model drift, and building a rollback plan before patient safety is at risk.
AI Risk & Governance
Who Owns the Alert? Building AI Escalation Pathways in the ED
AI alerts do not make the ED safer by themselves. A practical framework for assigning ownership, escalation, and off-ramps before go-live.