AI Risk & Governance
AI Change Control in the ED: Stop Model Drift Before It Reaches the Bedside

Why this matters
A practical ED playbook for tracking AI changes, detecting model drift, and building a rollback plan before patient safety is at risk.
Recommended next step
Pair this article with the free guide or course store if you want a more structured framework you can apply at the bedside or in leadership conversations.
The most dangerous AI problem in the emergency department may not be a spectacularly wrong answer. It may be a quiet change that nobody notices.
A triage model gets retrained. A vendor adjusts a threshold. The EHR team changes which lab value is sent to the service. A new patient population arrives. The alert still appears in the same place, with the same reassuring color, but it no longer means quite what clinicians think it means.
That is model drift. And unlike a dramatic outage, drift can look like normal work while it slowly changes the risk profile of a shift.
Emergency departments understand change control for medications, protocols, and equipment. AI needs the same discipline. If a tool can influence who gets roomed, worked up, or escalated, “the vendor updated it” is not a governance plan.
AI does not stay put after go-live
A pilot creates the illusion that an AI tool is a finished product. The team validates it against a defined population, watches a few dashboards, and then moves on to the next project. But the clinical environment keeps moving.
Staffing changes. Documentation habits change. A new lab instrument changes values at the edges. Boarding changes the time between triage and reassessment. Seasonal disease patterns shift. The model may also change as its developer updates code, weights, prompts, or reference data.
The result is not always a sudden drop in accuracy. More often, the relationship between the output and bedside reality becomes less dependable. A prediction calibrated last winter may over-call risk this summer. A documentation assistant may perform well for one physician’s shorthand and invent details when a new group uses different phrasing.
The NIST AI Risk Management Framework treats ongoing measurement and management as part of the AI lifecycle, not a one-time approval step. That same physician-led perspective is explored in Why Emergency Physicians Make the Best AI Governance Leaders. That is the right mental model for emergency medicine: the launch is the start of surveillance, not the end of evaluation.
Start with a change register, not a policy binder
Create one simple change register for every clinical AI capability. It does not need to be a 40-page committee document; a spreadsheet or ticket queue is enough if it answers five questions:
- What changed?
- Who approved the change?
- Which patients or workflows could be affected?
- What evidence supports keeping it live?
- What is the rollback plan?
The register should cover more than vendor releases: data feeds, interface placement, alert thresholds, prompt templates, inclusion criteria, user permissions, and downstream workflows. A “minor” interface change can alter response time; a new ordering default can change who receives a test.
Give each tool an owner with authority, not just a name on an org chart. That person should be able to pause the tool, request a review, and convene users. If nobody can stop the system on a night shift, nobody truly owns it.
I would also assign a clinical steward: a physician, advanced practice clinician, nurse, or pharmacist who understands the actual work. The technical owner can explain what changed in the code. The clinical steward can explain what changed in the room.
Define drift in terms clinicians can see
“Monitor performance” is too vague to guide a busy ED. Define drift with measures connected to clinical work.
For a triage or deterioration model, watch calibration, false-negative cases, overrides, time to reassessment, and whether high-risk patients reach a queue that can respond. For an imaging or sepsis flag, watch meaningful findings, missed cases, time from alert to action, and alert burden by shift. For ambient documentation, review omissions, fabricated details, correction time, and whether clinicians sign notes they did not adequately verify.
Do not rely on a single accuracy number. A model can maintain an acceptable average while failing for a subgroup, night shift, language preference, or care setting with sparse documentation. Stratify by factors that matter locally, while protecting privacy and avoiding simplistic conclusions from small samples.
Pair technical metrics with a short case review. Each month, examine cases where the output was wrong, ignored, or acted on. Ask what the clinician knew that the model did not—and whether the interface made the right action easier or merely made the wrong action feel official.
Build the off-ramp before you need it
Every clinical AI tool needs a downtime and rollback plan that works at 2 a.m. without a committee meeting.
The off-ramp should specify the trigger, authorized decision-maker, disablement steps, alternative workflow, and communication channel. “Call IT” is not enough; IT may turn off a service but not decide whether a patient population needs manual review.
Use graduated responses: a warning can prompt chart review; a stronger signal can pause new enrollments; a serious event may require disablement and notification.
Write the manual fallback as if the AI never existed. If the department cannot triage, reassess, document, or escalate without it, the tool has created dependency rather than support. Run the fallback in a tabletop exercise, then a low-stakes drill. The goal is to discover what people assume the system is doing.
The FDA’s current Clinical Decision Support Software guidance is a useful reminder that the function, intended use, and ability of a clinician to independently review the basis for a recommendation matter when determining how software should be evaluated and governed.
Make change review part of the operating rhythm
Governance fails when it is treated as an annual ceremony. A better approach is a short recurring review tied to the department’s existing operations cadence.
Review releases, safety events, override themes, subgroup signals, and open actions. Keep the meeting focused on decisions: continue, adjust, restrict, pause, or retire. Invite users, not only purchasers.
Require a brief “show your work” packet for material changes: what changed, why, what data were tested, which groups were included, known limitations, and what will be watched after release. If a vendor cannot provide enough detail for clinical review, that is itself a governance finding.
Document the version in the clinical environment. A clinician should be able to tell which model or configuration produced an output after a safety event. Without traceability, the department cannot reconstruct what happened or learn from it.
This is where physician leadership matters. The emergency medicine AI consensus statement emphasizes that AI should enhance, not replace, clinical judgment and that governance should remain physician-led. Keep accountability close to patient care.
Treat frontline overrides as safety data
An override is not automatically resistance. It may be the highest-value signal your department has.
Make it easy to record why a clinician ignored or corrected an output. Offer structured reasons—wrong data, context, timing, clinical relevance, or unavailable workflow—plus a short free-text field. Review themes, not individuals. If people are punished for reporting mismatches, dashboards get cleaner and truth gets scarcer.
Close the loop. When feedback changes a threshold, interface, or training message, tell staff. When it does not change the system, explain why. Trust grows when reporting a problem leads to a decision rather than a black hole.
And keep the scope honest. A clinical AI tool is not a substitute for staffing or a functioning escalation pathway. If the department’s response capacity is saturated, a more sensitive alert may create the appearance of vigilance while increasing noise. Change control must evaluate the whole workflow, not just the model output.
Dr. Chet’s Take
I am less worried about an ED that says, “We are not ready for AI,” than an ED that says, “We approved it last year, so we are done.”
The question is whether your organization can notice when a model is no longer good enough for the job you gave it. That requires an owner who can stop it, clinicians who can challenge it, a measurable definition of drift, and a practiced fallback.
If you cannot name the last meaningful change, the approver, and the trigger that would make you pause it, you do not have operational AI governance yet. You have software in a clinical workflow.
Key Takeaways
- AI governance continues after launch; model, data, workflow, and vendor changes can all create drift.
- Maintain a change register with a named technical owner, clinical steward, evidence summary, and rollback plan.
- Define drift using clinical measures, subgroup checks, override patterns, and case review—not a single accuracy score.
- Practice the manual fallback and make the pause authority explicit before a safety event occurs.
- Treat frontline overrides as safety intelligence and close the feedback loop with the people doing the work.
FAQ
How often should an ED review a clinical AI tool?
Use a recurring review cadence tied to operational work, with an immediate review after a material release, safety event, data-feed change, or meaningful shift in performance. High-risk tools may need monthly review; lower-risk tools still need a documented owner and trigger-based reassessment.
What counts as a material change?
Any change that could affect the output, the population receiving it, the clinician’s interpretation, or the ability to respond. That includes model updates, thresholds, prompts, data sources, interface placement, inclusion criteria, and downstream routing.
Who should be allowed to pause an AI tool?
The authority should be explicit on every shift. It may be shared between a clinical leader and technical owner, with a low threshold for temporary pause when safety is uncertain. Define who is notified and how the manual workflow starts.
Is an override rate a sign that the model is failing?
Not necessarily. An override may reflect a clinician recognizing context the model cannot see, or reveal a systematic problem. The useful question is why overrides happen and whether the reasons cluster by patient group, shift, workflow, or version.
What is the smallest useful first step?
Create the change register and write the rollback procedure for one high-impact tool. If your team cannot complete those two tasks, that is valuable information about readiness before adding another AI capability.
If you're an emergency physician (or any clinician treating patients daily) trying to understand how AI will actually impact your clinical practice -- not just the hype -- I put together a free practical guide. You can download it here: AI in EM Survival Guide
Read more from Dr. Shermer on Medium.
Sources
Keep reading
Related reading and your next step.
Ready to go further? Move from this article into structured training, scenario-based rehearsal, and more physician-written guidance.
Course
Translate the article into a repeatable framework
Use the physician-led course when you want a structured framework for evaluating AI tools, protecting clinical judgment, and leading implementation decisions.
Simulation
Practice the decision path under pressure
Use EM-Sim when you want scenario-based repetition that turns article-level insight into physician-facing emergency-medicine reps.
Blog
Browse more articles
Explore the full blog for more on AI in emergency medicine, then head to the course and simulation pages when you want the structured next step.
Related Articles
AI in Emergency Medicine
Before Your ED Buys AI, Demand the Evidence Packet
A vendor demo is not evidence that an AI tool will work in your emergency department. Here is the evidence packet and local stop rule to require before go-live.
AI Risk & Governance
Low Risk Is Not No Risk: Using AI Safely in ED Disposition
AI risk scores can help emergency physicians see patterns sooner, but low risk is not a disposition plan.
AI Risk & Governance
Who Owns the Alert? Building AI Escalation Pathways in the ED
AI alerts do not make the ED safer by themselves. A practical framework for assigning ownership, escalation, and off-ramps before go-live.