AI Risk & Governance

AI Change Control in the ED: Stop Model Drift Before It Reaches the Bedside

Chester Shermer, MD, FACEP August 16, 2026
AI Change Control in the ED: Stop Model Drift Before It Reaches the Bedside

Why this matters

A practical ED playbook for tracking AI changes, detecting model drift, and building a rollback plan before patient safety is at risk.

Recommended next step

Pair this article with the free guide or course store if you want a more structured framework you can apply at the bedside or in leadership conversations.

The most dangerous AI problem in the emergency department may not be a spectacularly wrong answer. It may be a quiet change that nobody notices.

A triage model gets retrained. A vendor adjusts a threshold. The EHR team changes which lab value is sent to the service. A new patient population arrives. The alert still appears in the same place, with the same reassuring color, but it no longer means quite what clinicians think it means.

That is model drift. And unlike a dramatic outage, drift can look like normal work while it slowly changes the risk profile of a shift.

Emergency departments understand change control for medications, protocols, and equipment. AI needs the same discipline. If a tool can influence who gets roomed, worked up, or escalated, “the vendor updated it” is not a governance plan.

AI does not stay put after go-live

A pilot creates the illusion that an AI tool is a finished product. The team validates it against a defined population, watches a few dashboards, and then moves on to the next project. But the clinical environment keeps moving.

Staffing changes. Documentation habits change. A new lab instrument changes values at the edges. Boarding changes the time between triage and reassessment. Seasonal disease patterns shift. The model may also change as its developer updates code, weights, prompts, or reference data.

The result is not always a sudden drop in accuracy. More often, the relationship between the output and bedside reality becomes less dependable. A prediction calibrated last winter may over-call risk this summer. A documentation assistant may perform well for one physician’s shorthand and invent details when a new group uses different phrasing.

The NIST AI Risk Management Framework treats ongoing measurement and management as part of the AI lifecycle, not a one-time approval step. That same physician-led perspective is explored in Why Emergency Physicians Make the Best AI Governance Leaders. That is the right mental model for emergency medicine: the launch is the start of surveillance, not the end of evaluation.

Start with a change register, not a policy binder

Create one simple change register for every clinical AI capability. It does not need to be a 40-page committee document; a spreadsheet or ticket queue is enough if it answers five questions:

  1. What changed?
  2. Who approved the change?
  3. Which patients or workflows could be affected?
  4. What evidence supports keeping it live?
  5. What is the rollback plan?

The register should cover more than vendor releases: data feeds, interface placement, alert thresholds, prompt templates, inclusion criteria, user permissions, and downstream workflows. A “minor” interface change can alter response time; a new ordering default can change who receives a test.

Give each tool an owner with authority, not just a name on an org chart. That person should be able to pause the tool, request a review, and convene users. If nobody can stop the system on a night shift, nobody truly owns it.

I would also assign a clinical steward: a physician, advanced practice clinician, nurse, or pharmacist who understands the actual work. The technical owner can explain what changed in the code. The clinical steward can explain what changed in the room.

Define drift in terms clinicians can see

“Monitor performance” is too vague to guide a busy ED. Define drift with measures connected to clinical work.

For a triage or deterioration model, watch calibration, false-negative cases, overrides, time to reassessment, and whether high-risk patients reach a queue that can respond. For an imaging or sepsis flag, watch meaningful findings, missed cases, time from alert to action, and alert burden by shift. For ambient documentation, review omissions, fabricated details, correction time, and whether clinicians sign notes they did not adequately verify.

Do not rely on a single accuracy number. A model can maintain an acceptable average while failing for a subgroup, night shift, language preference, or care setting with sparse documentation. Stratify by factors that matter locally, while protecting privacy and avoiding simplistic conclusions from small samples.

Pair technical metrics with a short case review. Each month, examine cases where the output was wrong, ignored, or acted on. Ask what the clinician knew that the model did not—and whether the interface made the right action easier or merely made the wrong action feel official.

Build the off-ramp before you need it

Every clinical AI tool needs a downtime and rollback plan that works at 2 a.m. without a committee meeting.

The off-ramp should specify the trigger, authorized decision-maker, disablement steps, alternative workflow, and communication channel. “Call IT” is not enough; IT may turn off a service but not decide whether a patient population needs manual review.

Use graduated responses: a warning can prompt chart review; a stronger signal can pause new enrollments; a serious event may require disablement and notification.

Write the manual fallback as if the AI never existed. If the department cannot triage, reassess, document, or escalate without it, the tool has created dependency rather than support. Run the fallback in a tabletop exercise, then a low-stakes drill. The goal is to discover what people assume the system is doing.

The FDA’s current Clinical Decision Support Software guidance is a useful reminder that the function, intended use, and ability of a clinician to independently review the basis for a recommendation matter when determining how software should be evaluated and governed.

Make change review part of the operating rhythm

Governance fails when it is treated as an annual ceremony. A better approach is a short recurring review tied to the department’s existing operations cadence.

Review releases, safety events, override themes, subgroup signals, and open actions. Keep the meeting focused on decisions: continue, adjust, restrict, pause, or retire. Invite users, not only purchasers.

Require a brief “show your work” packet for material changes: what changed, why, what data were tested, which groups were included, known limitations, and what will be watched after release. If a vendor cannot provide enough detail for clinical review, that is itself a governance finding.

Document the version in the clinical environment. A clinician should be able to tell which model or configuration produced an output after a safety event. Without traceability, the department cannot reconstruct what happened or learn from it.

This is where physician leadership matters. The emergency medicine AI consensus statement emphasizes that AI should enhance, not replace, clinical judgment and that governance should remain physician-led. Keep accountability close to patient care.

Treat frontline overrides as safety data

An override is not automatically resistance. It may be the highest-value signal your department has.

Make it easy to record why a clinician ignored or corrected an output. Offer structured reasons—wrong data, context, timing, clinical relevance, or unavailable workflow—plus a short free-text field. Review themes, not individuals. If people are punished for reporting mismatches, dashboards get cleaner and truth gets scarcer.

Close the loop. When feedback changes a threshold, interface, or training message, tell staff. When it does not change the system, explain why. Trust grows when reporting a problem leads to a decision rather than a black hole.

And keep the scope honest. A clinical AI tool is not a substitute for staffing or a functioning escalation pathway. If the department’s response capacity is saturated, a more sensitive alert may create the appearance of vigilance while increasing noise. Change control must evaluate the whole workflow, not just the model output.

Dr. Chet’s Take

I am less worried about an ED that says, “We are not ready for AI,” than an ED that says, “We approved it last year, so we are done.”

The question is whether your organization can notice when a model is no longer good enough for the job you gave it. That requires an owner who can stop it, clinicians who can challenge it, a measurable definition of drift, and a practiced fallback.

If you cannot name the last meaningful change, the approver, and the trigger that would make you pause it, you do not have operational AI governance yet. You have software in a clinical workflow.

Key Takeaways

  • AI governance continues after launch; model, data, workflow, and vendor changes can all create drift.
  • Maintain a change register with a named technical owner, clinical steward, evidence summary, and rollback plan.
  • Define drift using clinical measures, subgroup checks, override patterns, and case review—not a single accuracy score.
  • Practice the manual fallback and make the pause authority explicit before a safety event occurs.
  • Treat frontline overrides as safety intelligence and close the feedback loop with the people doing the work.

FAQ

How often should an ED review a clinical AI tool?

Use a recurring review cadence tied to operational work, with an immediate review after a material release, safety event, data-feed change, or meaningful shift in performance. High-risk tools may need monthly review; lower-risk tools still need a documented owner and trigger-based reassessment.

What counts as a material change?

Any change that could affect the output, the population receiving it, the clinician’s interpretation, or the ability to respond. That includes model updates, thresholds, prompts, data sources, interface placement, inclusion criteria, and downstream routing.

Who should be allowed to pause an AI tool?

The authority should be explicit on every shift. It may be shared between a clinical leader and technical owner, with a low threshold for temporary pause when safety is uncertain. Define who is notified and how the manual workflow starts.

Is an override rate a sign that the model is failing?

Not necessarily. An override may reflect a clinician recognizing context the model cannot see, or reveal a systematic problem. The useful question is why overrides happen and whether the reasons cluster by patient group, shift, workflow, or version.

What is the smallest useful first step?

Create the change register and write the rollback procedure for one high-impact tool. If your team cannot complete those two tasks, that is valuable information about readiness before adding another AI capability.

If you're an emergency physician (or any clinician treating patients daily) trying to understand how AI will actually impact your clinical practice -- not just the hype -- I put together a free practical guide. You can download it here: AI in EM Survival Guide

Read more from Dr. Shermer on Medium.

Sources

Keep reading

Related reading and your next step.

Ready to go further? Move from this article into structured training, scenario-based rehearsal, and more physician-written guidance.

Course

Translate the article into a repeatable framework

Use the physician-led course when you want a structured framework for evaluating AI tools, protecting clinical judgment, and leading implementation decisions.

Simulation

Practice the decision path under pressure

Use EM-Sim when you want scenario-based repetition that turns article-level insight into physician-facing emergency-medicine reps.

Blog

Browse more articles

Explore the full blog for more on AI in emergency medicine, then head to the course and simulation pages when you want the structured next step.

By using this site you agree to our Privacy Policy. We use cookies to keep you signed in. We do not sell your data.