Related on The Signal Room: Healthcare Leadership and Operational Reality. Related from HDSC: Why Healthcare Data Strategy Fails Without Operational Alignment.A major academic medical center utilizes a sepsis prediction model based on five years of archives of electronic health records. In the initial quarter, the model operates with optimal efficiency. In the eighth month, the false positive rates rapidly increase. Clinical staff conduct an investigation. They determined that a change in EHR documentation template occurred three months after model deployment. The change impacted the nursing assessment documentation. The model accessed a field that had a clinically different meaning from EHR training purposes. The documentation change was undocumented when it occurred, and the data science team was never informed. There was a lack of monitoring processes to track such a drastic change.
This situation is typical and, unfortunately, rarely publicly addressed. There is an expectation that AI models in healthcare analyze past data, implemented in real-time, and that the models will perform well as the documentation standards, clinical coordination, and patient demographics evolve. The failure to recognize models that have been validated as an ongoing form of healthcare AI oversight is an ongoing, unmitigated, and unaddressed healthcare AI issue.
The Model Drift Issue
Model drift occurs when the statistical correlation between input variables and outcomes changes as a result of deployment. In healthcare, it can occur for reasons that have no relation to the model itself. When a clinician changes within a practice, the clinician could have a different style of recordkeeping. Changes in the payer mix of a practice can occur. Surges in acuity can create a different clinical profile among the population that is being evaluated. The model will predictably create outcomes, but those outcomes are just a reflection of a reality that no longer exists.
An article in Nature Medicine describes the phenomenon of performance degradation in clinical prediction models and the measurable decline that occurs in the first twelve months after deployment. The article comments on the absence of institutional procedures to identify and/or address the decline and that performance of the models remains unmeasured and presumed.
The American Medical Informatics Association has addressed the need for the ongoing monitoring of models in clinical practice. Their research suggests that operational support requirements for enduring performance for AI models post implementation are grossly underestimated, resulting in models with excellent validation metrics that go on to degrade without being monitored.
Real World Evidence and Clinical Context
The concept of real world evidence, as defined by regulatory direction spanning from the 21st Century Cures Act to newer FDA frameworks, can be applied to AI model lifecycle management. Real world evidence focuses on understanding the impact of an intervention in practice and, therefore, involves the actual clinical setting, the real patient population, and the operational circumstances. The same applies to AI: a model can excel in a validation study (i.e., in the "lab") but may not perform similarly in the unpredictable world of clinical practice.
In response, the FDA highlighted these challenges in its draft policy on software as a medical device using AI/ML. The policy introduces a concept of a predetermined change control plan, which provides a framework of the types of changes a model can undergo post-market with descriptions on how those changes can be validated. This is a recognition that AI in healthcare is not a static product and that ongoing responsible change management is required.
Health Affairs has published research on the validation of AI models in practice and the gap in the health systems. The research highlights instances in which models that were trained at one institution and practice area were subsequently implemented at another institution and practice area and performed quite poorly, even with the same clinical use case. The variability in documentation, patient profiles, and care delivery across institutions introduces confounding factors that pre-deployment validation may not fully address.
Organizational Readiness
Keeping AI grounded in the real world requires organizational investment that extends well past initial deployment. The first of those is a model performance metrics monitoring infrastructure that is designed to check performance at regular intervals and to trigger model review when performance falls outside expected ranges. The second is a clinical user feedback mechanism that provides ongoing connection to the technical model maintenance teams. Clinicians are typically the first to detect when a model is producing outputs that are not clinically relevant and when this occurs, there should be a pathway to formalize these concerns.
The Office of the National Coordinator for Health IT has introduced interoperability standards that promote the data infrastructure that sustains model evaluation. In the absence of consistent, standardized data streams, monitoring is labor-intensive and sporadic, which is inadequate for models functioning in high acuity clinical settings.
Separately, the Regenstrief Institute has conducted research on the organizational aspects that lead to the successful deployment of AI in health systems over an extended period. Their research highlights the importance of having model operations, clearly defined escalation routes, and executive sponsorship. Institutions that view AI deployment initiatives as a singular project rather than an ongoing operational process consistently struggle to maintain the accuracy of the deployed AI model.
The Cost of Inattention
When a model drifts and no one is monitoring it, the consequences are not equally distributed. Patients in the most underrepresented segments in the training data suffer first the consequences of poor predictions. The clinical teams become frustrated and begin to circumvent the original workflow that the model was designed to support. For operational leaders, the technology investment has failed to deliver on the anticipated economic value. Collectively, the result is a loss of trust in the AI system, which makes the implementation of additional AI systems even more challenging.
For health system leaders, the discipline of keeping AI grounded is the discipline of treating deployed models as living systems that require continuous attention. The real question is, has the organization built the capacity to do this, or is it relying on the assumption that what worked at launch will continue to work indefinitely?
Context and Sources
Degradation of predictive clinical models has been examined in Nature Medicine. Continuous model monitoring in clinical settings has been covered by AMIA. The 21st Century Cures Act has put in place the real world evidence framework. The FDA has described the framework for AI/ML-based software as medical devices with predetermined change control plans. Health Affairs has reported on the validation vs real world AI performance gap. ONC has published interoperability frameworks that support the underlying data infrastructure. The Regenstrief Institute has studied the organizational issues related to the sustained use of AI. The current edition relates to operational and design aspects of Editions D, N, and V of this newsletter.
Christopher Hutchins
Founder & CEO, Hutchins Data Strategy Consultants