AI in Emergency Medicine
Before Your ED Buys AI, Demand the Evidence Packet

Why this matters
A vendor demo is not evidence that an AI tool will work in your emergency department. Here is the evidence packet and local stop rule to require before go-live.
Recommended next step
Pair this article with the free guide or course store if you want a more structured framework you can apply at the bedside or in leadership conversations.
At 0215, the vendor dashboard looks persuasive: a clean interface, a confident risk score, and a slide deck full of performance numbers. The overnight attending asks the question that matters: “Which of those patients look like my patients, on this shift?” Nobody can answer quickly.
That is the moment to stop. A hospital can buy software in an afternoon. It cannot buy evidence that a prediction is safe for its own patients. Before an AI tool influences triage, workup, escalation, documentation, or disposition, the emergency department needs an evidence packet it can read, challenge, and test locally.
The standard is “show us what this tool does, what data it needs, where it was tested, how it fails, and who can pause it.”
A vendor demo is not evidence for your ED
A demonstration answers one question: can the product produce an output? It does not answer whether that output is valid for your population, staffing pattern, or time-pressured decisions.
The FDA’s 2026 Clinical Decision Support Software guidance says that clinicians need enough information to independently review the basis of a recommendation. That includes the intended use, intended user, patient population, required inputs, data-quality requirements, algorithm-development approach, validation results, limitations, and patient-specific information used in the output. A score without that context is a conclusion without a chart.
The ONC HTI-1 final rule takes a similar approach for predictive decision-support interventions in certified health IT, giving users baseline information to assess fairness, appropriateness, validity, effectiveness, and safety. Turn that principle into a demand: no evidence packet, no clinical go-live.
The five pages every ED should request
Do not ask for “more transparency” and leave the request open-ended. Ask for five pages and give them to the people who will explain what changes on a real shift.
1. The intended-use page. Require the decision, users, eligible population, exclusions, and decisions the tool is not designed to make. Ask whether the output is advisory, interruptive, or automatically routed. The FDA guidance calls for clear identification of intended use, user, population, and whether the task is time-critical. If “risk stratification” quietly expands to “disposition support,” treat it as a new use.
2. The data page. Require every input, including vital signs, laboratory values, medications, free text, diagnosis codes, timestamps, and missingness rules. Ask when values are captured, how stale data are handled, and what happens when an input is absent or corrupted. The FDA guidance says users should identify required information, missing inputs, and patient-specific information used. A tool that cannot tell you what it did not see is not ready to guide a clinician.
3. The validation page. Request development and validation sites, dates, sample sizes, outcome definitions, subgroup results, and whether validation data were independent from training data. Ask for calibration, sensitivity, specificity, error counts, and confidence intervals when appropriate. The FDA guidance recommends enough detail to judge population fit and subgroup performance. The NIST AI Risk Management Framework calls for pre-deployment and regular in-operation testing, documented generalizability limits, and production monitoring.
4. The workflow page. Require the output’s location and timing, who sees it, the expected action, and what happens when no one acts. Ask what clicks, alerts, queues, or review tasks it creates. An accurate prediction that arrives too late or routes to an ownerless queue is not an effective intervention. Describe the manual workflow for downtime.
Use a short EM-Sim rehearsal before launch to test whether the right person notices, interprets, challenges, and acts when an output conflicts with bedside findings.
5. The change-and-exit page. Require the version, release process, notice for material changes, monitoring plan, incident route, pause authority, and rollback method. Ask what can change without approval: weights, thresholds, prompts, reference data, input mappings, or display logic. The NIST framework treats governance, measurement, and management as lifecycle work, including monitoring, feedback, incident response, recovery, and safe deactivation.
“The vendor monitors performance” is not a local safety plan. Name who can pause the tool at 2 a.m., the review trigger, the clinical alternative, and the time limit for restoring or retiring it.
Test the packet against your patients, not the sales deck
When feasible, start with a silent period: generate the output without displaying it, then compare it with what happened. Predefine the questions: did the tool receive the promised inputs, were values missing at the same rate, did the patient mix fit the validation population, and did the output arrive in time to matter?
Stratify results by locally meaningful groups and settings, such as age, sex, language, arrival mode, acuity, site, shift, and documentation completeness. Treat a small sample as a signal for further review, not proof of fairness or unfairness.
A 2025 prognostic study across seven Toronto hospitals found shifts tied to age, admission source, hospital type, and laboratory testing in 143,049 adult inpatients; some were associated with lower subgroup discrimination. Its monitoring and drift-triggered updating improved performance there. Read Detecting and Remediating Harmful Data Shifts in Deployment of Clinical AI Models as a reason to monitor, not copy without local judgment.
Outcomes are not a clean scoreboard. A clinician may see an alert, intervene, and change the outcome, making the model look more or less accurate depending on the measure. A 2025 NEJM AI analysis of postmarket surveillance describes this problem and argues that assessment must account for treatment effects and causal pathways. Track discrimination alongside overrides, time to action, missed cases, workload, and response capacity.
Put a stop rule in the contract
Every AI purchase needs a clinical stop rule before the first live output. Name safety events, pause authority, the manual process, documentation, and restart evidence.
The FDA guidance describes automation bias as over-reliance on an automated suggestion and notes that urgent decisions leave less time for review. A recommendation can be explainable yet unsafe if the clinician cannot review its basis before acting.
Write contract language that protects independent judgment. Require notice before material changes, version history, input definitions, a clinical incident contact, and export of outputs and supporting inputs. If a clinician cannot challenge the recommendation, that missing information is a procurement finding.
Dr. Shermer makes the same point in How Running a Telehealth Network Prepared Me for the AI Integration Challenge: the technology is not the deployment. Protocols, training, escalation, quality review, and refinement are the deployment. A contract covering uptime but not clinical integration is incomplete.
Dr. Chet's Take
I have spent 25 years in emergency medicine, and I have watched technology arrive with a better presentation than a plan. The first question is what it will change at the bedside. If the answer is “we are still working that out,” you are ready for a test, not a live clinical path. That distinction protects patients and clinicians.
That being said, I do not want emergency physicians standing outside every AI project with their arms crossed. We should be in the room early. Ask for the inputs, validation, missing-data behavior, and pause authority. HEMS and National Guard work taught me that a system is only as safe as its fallback. A polished dashboard is not one. A practiced manual workflow is.
If you are leading an emergency department, add an evidence-packet checklist to procurement and run one case through it with the people who will use the tool. Do not wait for a safety event to discover that nobody owns the pause button. Clinical AI earns its place by surviving scrutiny before go-live.
Key Takeaways
- A vendor demonstration proves that software can produce an output; it does not establish validity for your emergency department.
- Require five pages before go-live: intended use, data inputs, validation, workflow, and change-and-exit rules, using the FDA’s independent-review principles.
- Test locally and monitor drift because performance can change when patients, sites, workflows, or measurements change (JAMA Network Open).
- Put pause authority, the manual fallback, version history, incident route, and restart criteria in writing.
- If clinicians cannot understand what the tool saw and when to reject it, it is not ready for an ED workflow.
FAQ
What should an emergency department ask an AI vendor before buying?
Ask for intended use, population, exclusions, inputs, missing-data behavior, validation, subgroup performance, workflow, version history, monitoring, and rollback. The FDA CDS guidance provides a starting point.
How can an ED validate an AI tool locally?
When feasible, use a silent period, compare outputs with real cases, and review timing, misses, overrides, and response capacity. Treat small samples as signals, not proof.
Who should be allowed to pause an AI tool?
Name the clinical stop authority before deployment and make it available every shift. The pause process should start the manual workflow, preserve outputs and inputs, notify owners, and define restart evidence; the NIST AI Risk Management Framework includes monitoring, incident response, recovery, and deactivation.
Does FDA guidance mean every clinical AI tool is a medical device?
No single label answers that question for every function. The FDA guidance discusses independent review, automation, and whether the task is time-critical. Assess the specific function and intended use.
If you're an emergency physician (or any clinician treating patients daily) trying to understand how AI will actually impact your clinical practice — not just the hype — I put together a free practical guide. You can download it here: AI in EM Survival Guide.
Sources
- U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff
- Office of the National Coordinator for Health Information Technology. HTI-1 Final Rule
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- Subasri V, et al. Detecting and Remediating Harmful Data Shifts in Deployment of Clinical AI Models
- Challenges in the Postmarket Surveillance of Clinical Prediction Models
- Dr. Chet Shermer. How Running a Telehealth Network Prepared Me for the AI Integration Challenge
- Related simulation training: EM-Sim
- Books by Dr. Shermer
Keep reading
Related reading and your next step.
Ready to go further? Move from this article into structured training, scenario-based rehearsal, and more physician-written guidance.
Course
Translate the article into a repeatable framework
Use the physician-led course when you want a structured framework for evaluating AI tools, protecting clinical judgment, and leading implementation decisions.
Simulation
Practice the decision path under pressure
Use EM-Sim when you want scenario-based repetition that turns article-level insight into physician-facing emergency-medicine reps.
Blog
Browse more articles
Explore the full blog for more on AI in emergency medicine, then head to the course and simulation pages when you want the structured next step.
Related Articles
AI Risk & Governance
AI Change Control in the ED: Stop Model Drift Before It Reaches the Bedside
A practical ED playbook for tracking AI changes, detecting model drift, and building a rollback plan before patient safety is at risk.
AI Risk & Governance
Low Risk Is Not No Risk: Using AI Safely in ED Disposition
AI risk scores can help emergency physicians see patterns sooner, but low risk is not a disposition plan.
AI Risk & Governance
Who Owns the Alert? Building AI Escalation Pathways in the ED
AI alerts do not make the ED safer by themselves. A practical framework for assigning ownership, escalation, and off-ramps before go-live.