An average health client has begun its third use case for artificial intelligence in the most recent eighteen months. The first was a model of risks for readmissions. The second was for clinical documentation improvements. The most recent is an engine for optimizing patient scheduling. Each one was approved through a separate process, handled by separate teams, and was documented in separate formats. The Chief Medical Officer has asked for an overview of the current AI tools employed, the various review processes, performance metrics, etc. Currently no one possesses the capacity to consolidate that data. The health system has a range of artificial intelligence tools and none of the systems to manage AI in a production environment at scale.
Early adopter health systems experience the same thing. The first deployments of artificial intelligence were all seen as separate and distinct. Each had its own distinct champion, different approval processes, and heterogeneous performance metrics. The deficit was the lack of standard processes that were flexible enough to adapt to the use case without rebuilding the system. The most visible and most expensive absence is the lack of that structure.
Limitations of Project-Specific Oversight
When it comes to oversight of artificial intelligence, if it is limited to individual projects, it creates use cases ad infinitum, and every new use case feels like reinvention of the wheel. From the clinical, legal, compliance, IT, and data science perspectives, the engagement is case by case. If there are only one or two models in production, the situation is manageable. However, given current estimates, active artificial intelligence tools are projected to exceed ten systems, and the situation becomes untenable.
The Journal of the American Medical Informatics Association published a study on the Operational Burden of the Project-Based AI Review in Health Systems. The study described ad hoc review processes causing health systems to experience longer deployment times, reviews are prone to duplication of efforts, and there is a lack of consistency in documenting model performance and risk assessment. The study specifically states that project-based oversight is not sustainable.
Similarly, the American Hospital Association has noted that health systems are in need of transitioning from project-based AI systems to enterprise-based AI systems. Their review identified lack of pre-established criteria for review, performance tracking systems, and cross-team collaboration as the main reasons, among many, that responsible AI scaling is hindered.
What a Scalable Structure Requires
A scalable oversight framework has four traits. Firstly, it has uniform techniques for assessment across the various clinical and operational domains involved with the artificial intelligence (AI) use cases. For example, a readmission model and an appointment scheduling model would be assessed using the same structure, with consideration for the risk and clinical impact, rather than completely different assessments.
Moreover, it integrates the performance metrics of any existing AI tools. KLAS research has shown that many healthcare institutions cannot cite all operational AI tools within their enterprise, and without unified oversight, it can only ever be reactive and incomplete.
Finally, it mandates certain roles that entail authority. The Office of the National Coordinator for Health IT (ONC) recommends that health systems designate individuals or teams as having explicit authority concerning the entire lifecycle of the AI model, including pre-deployment assessment, active surveillance, and model retirement. These roles require permanent staff commitments, enough funding, and clear authority from the institution.
Fourth, it uses a layered risk classification system. The FDA has defined risk categories for artificial intelligence and machine learning combined with software as a medical device. Health systems could use a similar system, internally matching AI tools with defined clinical risks and applying control measures proportional to that risk. An operational tool that is low risk should not be controlled as if it were a high-risk diagnostic model. Rate-based control for both tools will result in unnecessary delays and lack any meaningful safety improvements.
The Coordination Problem
A certain degree of scalable oversight will entail a level of grading, so it will require coordination across functions that have typically been siloed. Clinical leadership, IT, Data Science, law, compliance, and quality improvement all have a role in overseeing AI, and most health systems have these groups working in parallel instead of being integrated.
The National Academy of Medicine has outlined such integrated coordination with respect to AI in healthcare. They point to integrated decision-making, inconsistent terminology across disciplines, and disparate objectives as a persistent impediment to effectively scalable AI oversight.
Health Affairs researches health systems with any form of enterprise-level AI oversight, and the most successful systems have the following: an AI standing committee composed of clinical, technical, legal, and operational leadership; a unified data management framework for assessing AI performance; and a clearly articulated concern-reporting process, with specified time frames for issues to be escalated to the executive level.
From Reactive to Structural
The transition from project-driven to enterprise level AI management captures how fast health systems can innovate in the health sector and is a landmark achievement for health systems. There is a market opportunity for early adopters of enterprise level AI stewardship to operationalize use cases in their systems swiftly. They will be able to proactively identify and eliminate operational issues that may risk patient safety. They will also be able to operationalize their AI initiatives and oversight to a degree that will surpass the reliability of existing clinical performance initiatives. This will be a risk mitigation structure for the public, the regulators, and the payers. They will demonstrate to the public, the regulators, and the payers that their AI initiatives are controlled to the degree that it is a risk mitigation structure for the performance of other clinical initiatives.
Institutions that defer this transformation will be unable to cope with the growing number of AI tools and will rely on processes that were designed for a much smaller set of tools. The clinical risk, operational inefficiencies and institutional reputation will determine the impact of the lack of change.
Sources and Context
JAMIA has been publishing its findings concerning operational issues with the project-based AI review. The American Hospital Association speaks to the concerns with the project-based AI to enterprise AI management. KLAS has documented the inequities in inventory related to AI initiatives in the health systems. ONC has proposed specific responsibilities related to AI model lifecycle management. The FDA has proposed to risk categorize AI/ML software as medical devices. The National Academy of Medicine has addressed the necessity of collaboration in multiple disciplines concerning AI. Health Affairs has assessed the enterprise level AI oversight models in health systems. There is resonance concerning operational readiness and institutional design in editions D, W, and X of this newsletter.
Christopher Hutchins
Founder & CEO, Hutchins Data Strategy Consultants