Companion reading: Best Healthcare AI Newsletters for Executives, An Honest List: https://aihealthpulse.beehiiv.com/p/best-healthcare-ai-newsletters-for-executives-an-honest-listWas this edition forwarded to you?
AI Health Pulse is a weekly briefing on healthcare AI strategy and oversight, written for health system leaders. Get every edition in your inbox.
Subscribe at https://aihealthpulse.beehiiv.com/subscribe

In 2021, JAMA Internal Medicine published an article reviewing the external validation of a sepsis prediction model built to operate within electronic health records, which had been integrated into hundreds of hospitals. Reported performance metrics for the model suggested that the model was robust. However, when the model was applied to actual hospitalizations of a large academic health system, the model generated alerts for a large proportion of admitted patients, while the model was robust to most cases of sepsis. The model was robust to reported performance metrics because no one validated beyond the reported performance metrics.

The most troubling aspect of this case of external validation is that the model remained unchanged. What was altered was the context of the model evaluation. Validation metrics are designed around a single population and a single definition for the outcome for which the model would be applied. Move the model to another hospital with a different patient population and a different hospital culture, and the metrics may be completely meaningless.

What the Metric Cannot See

Accuracy metrics address a particular question: of the data the model encountered, how often did the model’s risk ranking correspond with the outcome? Everything else falls outside of this question (e.g., the timeliness of the alert, what actions are available to the clinician, etc.). If a model’s classification is inconsistent, then the model is being evaluated externally with a key answer that is also inconsistent. Discrimination and calibration are two different aspects of model performance. A model can be great at discrimination and poorly at calibration. Conversely, a model can be well calibrated but poorly defined if the system’s response to the risk is already implemented.

An executive has a minute to spend on the calibration and discrimination of a risk model. Discrimination is the model’s ability to rank sicker versus healthier patients. Calibration, on the other hand, is the model’s ability to maintain a 20% risk prediction that ultimately is reflected in 20% of the cases. Thus, a risk model can have good discrimination in ranking higher-risk patients, but poor calibration that is enough to cause risk of hitting the alert threshold. From a sales point of view, vendors prefer to focus on discrimination versus calibration.

The most intelligent deceptions scores use subtlety, and imprecision acts as the most effective tool in the subtle arsenal. Once a model begins to issue notifications that coincide with the triggered care team outcomes, the model begins to earn points for the coinciding predictions. The care team begins to work on sepsis responses after a model sends a notification. In this example, the model begins to earn points for confirming clinicians’ predictions. Even though care team members work on sepsis responses after the notification, the model earns points for making a correct prediction. In this example and other similar situations, the prediction seemingly maintains the status quo of clinical practice. The model earns no distinction in issuing notifications before or after the care team completed the outcome.

Not a single example demonstrates this issue. A demonstration shows a perfect metric, and the model earns that metric. At some point with a perfect score, the model begins to earn points for operational value on the buyer’s dollar. In my years as a data leader, I frequently witnessed a model that earned a perfect score where no one behaved differently during the actual workflow.

The Question the Demonstration Never Answers

Sepsis validation occurred after researchers demonstrated the deployed model used to evaluate a disparate dataset within a local context. This reflects a methodology, and one that all health systems have access to prior to procurement, as opposed to post-hoc. Request the vendor demonstrate the model’s performance on your patients, and evaluate outcome definitions as you would, and don’t accept “no.” An accuracy metric generated on a disparate dataset is a starting hypothesis, and not a justification for implementation within your health system.

Rather than including it in discussions, incorporate it in the contract. A vendor that has confidence in their model should be able to accommodate a feasibility study with local validation and the metrics defined by the customer as a pre-requisite to full settlement. They should also be fine with the metrics being repeated on a set cadence post go-live. If they have no confidence in the model, consider that their lack of confidence will likely be the case for you.

One of the many ways to safely evaluate a model is to allow it to be deployed in a “quiet” mode, in which it generates predictions on live, local data and does not notify anyone. For a set time period, e.g. three months, the model can be kept in quiet mode, and potentially the only cost of this would be the time of the analysts. Vendors are very used to this methodology and have to accommodate this as a stipulation of the pilot.

Cost conceals itself amidst noise. Well-measured models that fail to implement any meaningful changes generate waste and distraction. These distractions prompt clinicians to overlook alerts, which leads to a trend throughout the organization to ignore all alerts—even those that are genuinely beneficial. Given this situation, alerts do the opposite of making an organization safe, even if the measuring tool that generates the alerts has a good score.

This is the Real Work

This is the emphasis of one of the chapters in my book, Beneath the Signal. Here, I explore the reasons a measuring tool can ultimately let down the organization that implements it. Additionally, I explore the gap that exists between validation metrics and operational value, as well as the costs owed by the organization that stem from this gap. This chapter is available for free. To get a free chapter, go to https://hutchinsdatastrategy.com/free-chapter. You can also get this briefing each week, for free, by sending me an email.

Essentially, the true value of a system is found in the data pipelines that feed the model, the workflows that receive its outputs, and the people that decide to trust it, as well as the procedures that determine how the accuracy of a model’s outputs may deteriorate. Investing in the model doesn’t ensure any of the components mentioned above are included. If health systems see validation as a purchasing gate and nothing more, that testing has shown the least valuable interface. The interfaces that determine the value—data and workflows—are local in nature and cannot be validated from the outside.

Changes can be made at the inputs as well. For example, lab interfaces can be upgraded, documentation templates can be altered, a necessary field for the model can be submitted as empty, and the model can continue to score predictions as if nothing has changed.

Security develops in much the same way. My recent The Signal Room interview with Pranava Adduri covers how AI agents are the next major healthcare security risk. It makes the security argument that the threat lies in the surrounding system of the model, and there is no product benchmark to encapsulate it. You can listen to the episode here: https://signalroompodcast.com/episodes/ai-agents-healthcare-security-risk

What Should Leaders Demand

Before using your system on real patients who will produce real outcomes, demand local validation, and determine real outcomes. Before deployment, describe the operational metric beside the statistical one, such that the time to do something (e.g. modify the treatment) is well understood, and not a matter of guesswork. Also design for routine revalidation. Take note that a model's performance is not inherent and cannot be assumed. A model’s performance is a dynamic entity, and as the context and the target population of the model evolve, its performance will change. A performance metric assumed at the time of deployment will deteriorate. Performance monitoring is essential, and a named individual must be designated. Once a model is deployed, the expectation that performance will be monitored exists, even if it is not expressly stated. The sepsis model is helpful in understanding this. The model was in widespread use for some time before local validation of performance was carried out by independent researchers. In this instance, the information that validated local performance was accessible, yet no one had been designated to monitor it.

Two deployments of the same model can be made to function like different products. The alert threshold determines how many alerts are fired and how many are false alerts. The routing determines whether the alert actually reaches someone who can take action before it is too late. Both settings are made after the purchase, and they can invalidate a model that passes all of the tests.

This does not mean you need a data science department comparable to a payer. It means you need to make sure that a model will earn the trust of local evidence, and that this evidence will be provided on a regular basis. The cost of keeping the local evidence is far less than maintaining an alerting system that clinicians have forgotten about.

Context and Sources

This edition draws on the 2021 JAMA Internal Medicine external validation of a widely deployed sepsis prediction model and on the published distinction between discrimination and calibration in clinical prediction work. It adapts territory from Beneath the Signal, the free chapter of which is available at https://hutchinsdatastrategy.com/free-chapter, and continues themes from issue 44, The AI Readiness Diagnostic, issue 53, The Population the Model Cannot See, and issue 17, Trust by Design.

The AI Health Pulse is a weekly briefing on healthcare AI strategy and oversight that provides independent insights for busy executives. It is written by Christopher Hutchins, a former health system data executive and the founder of Hutchins Data Strategy Consultants. The publication is free from sponsorship and paid placements. Each edition is based on named sources, and the operational aspects of healthcare AI strategy and oversight are covered, such as data readiness, model oversight, and the impact of AI on an organization. You can subscribe at https://aihealthpulse.beehiiv.com/subscribe. The complete archive is also free to access.

Companion reading: Best Healthcare AI Newsletters for Executives, An Honest List: https://aihealthpulse.beehiiv.com/p/best-healthcare-ai-newsletters-for-executives-an-honest-list

Christopher Hutchins
Founder & CEO, Hutchins Data Strategy Consultants

Continue reading from Hutchins Data Strategy

On the Signal Room podcast

More from Chris Hutchins

Listen: The Signal Room, where Chris talks with the people building healthcare AI and the leaders who answer for it. On your podcast app or on YouTube.

Read: Beneath the Signal. One email at https://hutchinsdatastrategy.com/free-chapter gets you a free chapter plus The AI Health Pulse every week.

Subscribe: AI Health Pulse, every week at https://aihealthpulse.beehiiv.com/subscribe

Book Chris to Speak: https://www.chrisjhutchins.com/
Learn about our work to improve mental health around the world: https://continuamindhealth.com/
Do you have a podcast? We offer a Free growth assessment from PodcastPull.com: https://podcastpull.com/
Read on Substack: https://aihealthpulse.substack.com/

Tags: AI Health Pulse newsletter · clinical AI validation · model accuracy · healthcare AI value · sepsis prediction · calibration · Beneath the Signal

Recommended for you

View all
caret-right