Generative AI medical device being monitored by a physician

Generative AI Medical Devices Need a Post-Deployment Drift Budget

By Gleb Tsipursky, PhD

Generative artificial intelligence is creating a new class of challenges for medical-device developers, clinicians, regulators, and health systems. Unlike traditional software, generative AI systems can produce varied and open-ended outputs, making it harder to determine performance through a single premarket evaluation.

The U.S. Food and Drug Administration (FDA) is now examining these challenges directly. In August 2026, the FDA published a discussion paper on considerations for regulating generative AI-enabled medical devices. The paper addresses risk assessment, premarket evaluation, postmarket monitoring, foundation models, and agentic AI systems. The FDA emphasizes that the discussion paper is intended to gather stakeholder feedback and does not represent final guidance or regulatory requirements.

That shift toward lifecycle monitoring matters because a medical AI system does not stop creating operational risk after it passes an initial evaluation.

Real-world clinical environments are different from controlled testing environments. Patient populations change. Clinical workflows evolve. Documentation practices vary. Software components are updated. New edge cases appear. Clinicians also differ in how they interact with AI-generated outputs.

For that reason, health systems should consider establishing a post-deployment drift budget.

What Is a Post-Deployment Drift Budget?

Post-deployment monitoring of a generative AI medical device

A post-deployment drift budget is an operational framework for measuring how much human intervention a clinical AI workflow requires after deployment.

The idea is simple.

During the first 30 days after a consequential generative AI system enters clinical use, the organization should measure the human effort required to keep the workflow safe, accurate, and reliable.

That can include:

– Correcting AI-generated content
– Verifying clinical information
– Escalating unusual cases
– Rewriting patient-facing communications
– Rejecting inappropriate recommendations
– Checking sources or supporting evidence
– Recovering from failed outputs
– Correcting downstream documentation
– Deciding not to use the AI system in a particular case

The objective is not to eliminate human intervention.

In healthcare, human oversight can be essential to safe AI deployment. The objective is to understand whether the intervention burden is acceptable, whether recurring problems decline, and whether the workflow remains reliable as real-world conditions change.

The concept aligns with the FDA’s broader work on postmarket monitoring of AI-enabled medical devices. FDA research specifically examines methods for detecting changes in inputs, monitoring output performance, and understanding why AI performance may vary across populations, sites, and clinical conditions.

Establish the Clinical Baseline Before Deployment

Clinician comparing traditional and AI-assisted healthcare workflow

A health system cannot determine whether AI is improving a workflow without first understanding the existing workflow.

Before deployment, the organization should define the specific clinical task the system is intended to support.

That could include:

Clinical documentation
– Patient communication
– Medical image interpretation support
– Triage assistance
– Coding
– Clinical decision support
– Administrative workflows

The baseline should include more than model accuracy.

Organizations should measure how long the existing process takes, how often clinicians correct errors, how frequently cases require escalation, and what happens when an error reaches the next stage of the workflow.

This distinction is important.

A generative AI tool may reduce the time required to produce an initial output while increasing the amount of time clinicians spend reviewing and correcting that output.

If organizations measure only generation speed, they may conclude that the system is improving productivity when the total human workload has actually increased.

Count Human Interventions, Not Just AI Errors

The next step is to establish an intervention ledger.

Every meaningful intervention should be recorded during the initial deployment period.

For example, a clinician may need to:

1. Correct an inaccurate statement.
2. Remove unsupported information.
3. Resolve a contradiction in the medical record.
4. Rewrite a patient-facing explanation.
5. Verify a clinical recommendation.
6. Escalate a case to another clinician.
7. Recover after the AI fails to complete a task.
8. Reject the AI output entirely.

Not every intervention represents an AI failure.

A clinician reviewing a generated note may simply be performing an appropriate safety check. The more important question is whether the organization understands the type, frequency, and cost of those interventions.

A useful intervention ledger can record:

– Type of intervention
– Time required
Clinical significance
– Whether the issue has occurred previously
– Whether the issue was detected before reaching the patient
– Whether a workflow or technical change could prevent recurrence

Over time, this creates an evidence base for deciding whether the AI system is becoming easier or harder to operate safely.

Separate Benign Variation From Consequential Drift

Clinician reviewing consequential AI output changes

Generative AI naturally produces variation.

Two outputs generated from similar inputs may not use identical wording. That does not automatically indicate a safety problem.

A monitoring system should therefore distinguish between benign variation and consequential drift.

For example, a slightly different sentence structure in a routine administrative note may have little clinical significance.

A changed statement about:

– Medication
– Symptoms
– Follow-up instructions
– Clinical risk
– Diagnosis
– Patient instructions

may require substantially more scrutiny.

This distinction allows health systems to focus their monitoring resources on clinically meaningful changes rather than treating every output difference as a defect.

It also gives technology teams better information about which failure categories require additional safeguards, workflow redesign, model evaluation, or narrower intended use.

Measure Recovery Time

Clinician recovering a healthcare AI workflow after an incident

AI governance frequently focuses on preventing errors.

Healthcare organizations also need to measure what happens when prevention fails.

For every consequential incident, organizations should consider measuring the time from detection to a safe state and then to full workflow recovery.

Questions should include:

– Did the clinician know how to stop using the system?
– Could the team determine what information influenced the output?
– Could incorrect documentation be corrected?
– Could a patient message be withdrawn or corrected?
– Did someone have clear ownership of the incident?
– Did the organization know when the workflow could safely resume?

Recovery time matters because a relatively rare failure can still create significant risk if the organization cannot identify, contain, and correct it quickly.

The FDA’s work on real-world AI performance similarly recognizes the importance of monitoring changes and understanding performance variation after deployment.

Test the Workflow With a Second Clinician

A clinical AI workflow should not depend on the one clinician or implementation specialist who originally helped build it.

During the first 30 days, a health system should consider conducting a second-operator test.

A qualified clinician who was not involved in the original implementation should operate the workflow using the documented procedures.

That clinician should be able to:

– Understand the system’s intended use
– Review consequential outputs
– Identify situations requiring escalation
– Recognize unsafe or unreliable outputs
– Follow the documented recovery procedure

If the second clinician cannot safely operate the workflow without private coaching from the original implementation team, the organization has identified an important operational weakness.

The technology may function correctly, but the operating model is not yet transferable.

This test can expose hidden knowledge.

The original implementation team may know which prompts work best, which patient situations require caution, which output patterns deserve additional review, and which system behaviors indicate a potential problem.

If that knowledge exists only in people’s heads, the workflow remains fragile.

Set a Drift Threshold Before Scaling

The organization should establish its decision criteria before expanding the system.

For example, expansion might require:

– Stable or declining consequential intervention rates
– Acceptable recovery times
– Reduction in recurring failure categories
– Successful second-clinician testing
– No unresolved high-severity safety concerns
– Evidence that the workflow continues to deliver meaningful value

If the system creates value but generates excessive verification work, the organization may need to redesign the workflow.

If particular patient contexts repeatedly produce problems, the intended use may need to be narrowed.

If the human burden or clinical risk becomes unacceptable, the organization should be prepared to pause or stop the workflow.

The key is to establish these thresholds before enthusiasm for a promising technology makes objective decision-making more difficult.

Why Post-Deployment Monitoring Matters for Generative AI

The FDA’s 2026 discussion paper highlights a central challenge with generative AI-enabled medical devices: their outputs can be varied and open-ended, while the systems themselves may undergo changes after deployment. The FDA is therefore considering approaches such as periodic benchmarking and sample-based clinician review as possible postmarket monitoring strategies.

This is particularly relevant because AI performance can change when real-world inputs, patient populations, clinical environments, or software components change.

FDA research on AI-enabled medical devices has already identified data drift and changes in clinical conditions as important factors that can affect the safety and effectiveness of AI systems over time.

A post-deployment drift budget complements this broader lifecycle approach by adding an operational question:

How much human work is required to keep the AI-supported workflow safe and useful?

That question is important because technical performance metrics do not always capture the full burden placed on clinicians.

Healthcare AI Governance Needs Operational Evidence

The American Medical Association has also emphasized the importance of governance, physician involvement, transparency, data quality, cybersecurity, and monitoring as AI moves from experimentation toward broader clinical use.

The AMA’s current AI policies also emphasize physician involvement throughout the AI lifecycle, including development, governance, clinical integration, and postmarket surveillance.

These principles become more useful when organizations translate them into measurable operating practices.

A drift budget gives clinicians, technology teams, risk leaders, and executives a shared language.

Instead of asking whether an AI tool simply “works,” they can ask:

How often does it require correction?

It depends on the type of AI system and the clinical task. In the early stages, some outputs may need regular review or correction. The important thing is to track how often corrections are needed and whether that number decreases as the system and workflow improve.

What types of failures recur?

Common problems can include inaccurate information, unsupported statements, missing details, contradictions in medical records, or inappropriate recommendations. Tracking these recurring problems helps the team understand where the system needs better controls or workflow changes.

How much verification time does it create?

The AI may save time by creating an initial draft or recommendation, but clinicians still need to review important outputs. The organization should measure the actual time spent checking and correcting AI-generated content to determine whether the system is truly saving time overall.

How quickly can clinicians recover from failures?

Clinicians should be able to recognize a problem, stop using the AI output when necessary, correct the information, and safely continue the workflow. If recovery takes too long or nobody knows what to do, the organization needs a better recovery process.

Does performance change across patient populations or workflows?

Yes, it can. An AI system may perform differently with different patient groups, clinical settings, data sources, or workflows. This is why performance should be monitored in real-world use rather than relying only on the results from the original evaluation.

Can another qualified clinician operate the system safely?

Ideally, yes. A second clinician who was not involved in the original implementation should be able to understand the workflow, review the AI’s output, identify problems, and follow the safety procedures without needing private coaching from the original team.

Does the system continue to provide enough value to justify its operational burden?

That’s the key question. An AI system should not be considered successful simply because it produces results quickly. The organization should compare the benefits, such as time saved or improved workflow, with the human effort required for verification, correction, escalation, and recovery. If the burden becomes too high, the workflow may need to be redesigned or its use limited.

The Goal Is Not Zero Human Intervention

A mature healthcare AI strategy should not treat every human correction as evidence that the technology has failed.

Human judgment remains central to many clinical workflows.

The better objective is to understand when, why, and how much human intervention is required, and whether that burden remains acceptable as the system scales.

Generative AI medical devices will continue to evolve. Regulation will evolve with them. Health systems therefore need mechanisms that allow them to observe what happens after deployment, not just what happened during testing.

A post-deployment drift budget provides one possible framework.

By measuring interventions, distinguishing consequential drift from benign variation, tracking recovery time, testing workflow transferability, and establishing predefined scaling thresholds, healthcare organizations can build stronger evidence around the real-world performance of clinical AI.

The future of responsible medical AI will not be determined only by whether a model performs well in a controlled evaluation.

It will also depend on whether healthcare organizations can continuously demonstrate that the technology remains safe, useful, manageable, and worthy of clinical trust after it enters the real world.

About the Author

Gleb Tsipursky, PhD, is a behavioral scientist and CEO of Disaster Avoidance Experts. He is the author of The Psychology of AI Adoption at Work: From Resistance to Results, http://disasteravoidanceexperts.com/aibook published by Georgetown University Press. The book examines the human and organizational factors involved in responsible AI adoption, including trust, risk management, governance, and workforce adoption. 

Resources