There appear to be common patterns in the multiple AI applications in healthcare. Each new application has a clearly defined use case, a particular team, and a well-calibrated and well-prepared dataset. There is also insufficient openness regarding the level of oversight and editing in front of the outputs. Most of the time, the results are positive. The model works, the organization restructures the workflow, and the organization feels a sense of improvement.

The topic of discussion shifts to scaling, and this seems to be where the most work changes, even if it is very minor. This has been my experience in multiple settings, and I feel that the early wins are a result of how tightly controlled those few initial win conditions were, and not because of the few initial conditions in this closed loop system.

In contrast, scaling often introduces the model to a new system that has been evolving through a set of constraints. Purpose changes, systems modify, and workarounds are integrated as part of the workflow. This may not be well captured or represented in the data. The pilot indicates a system that has been refined through purposeful effort, and in other instances, because the team was very close to the work and were able to continuously improve that particular system.

Due to the pilot, the team has meticulously curated the data, adjusted definitions, and has built a model to the level where they can articulate the inputs and outputs, which is a credit to the pilot. However, when this model is applied more broadly, it will inevitably confront numerous longstanding issues, including the variability of documentation pertaining to the same case, as well as the usual contextual variations when definitions are deployed to data during the preparatory phase before it is deployed.

The primary focus of the work is the model. However, the outputs and value of the work done are largely a function of the ability of the model to distinguish and process the contextual variations and much less a result of the inherent capabilities of the model.

The distinction between showing a use case and simply assuming it becomes evident. A pilot demonstrates that a use case can function within a specific set of circumstances. It validates that the approach is feasible. However, when it comes to scaling, it becomes critical to ensure that those circumstances are applicable across environments that have evolved independently over time.

In the majority of organizations, this has not been the case.

Recognizing the data environment directly is much clearer. Analysts have to make data usable by reconciling, cleansing, and validating. That is the nature of the work, and it is expected. However, it is not reflected in the outputs nor in the descriptions of the work at the leadership level. It impacts the timelines, the level of confidence, and the frequency results must be explained before being accepted.

The system continues to work as usual.

The work of making the system usable is hidden in the costs that are required.

The same trend is noted in the deployment of AI.

Models expect the input data to come from a system that is structured enough to allow for interpretation. The variation of the same concept across the system leads the model to represent that variation, even if the system is in a contradiction. The output can be stable even when the input conditions are drastically different.

This is a common occurrence across different situations. It is observed in the research published in JAMA Network Open and Nature Medicine, that model performance discrepancy is due to the data environment and not the model across different institutions. Also, the NIST AI Risk Management Framework cites quality, representativeness, and context of data as core risk factors, and these factors are present in most organizations.

AI visibility manifests these factors.

For Chief Information Officers and Chief Data Officers, this line of work is more about experience than anything else. They define parameters across domains, fix data pipeline duplication, and ensure manual reconciliation is not required. Engage clinical and operational stakeholders to define data from source systems as it is generated.

This type of work is often slow to fund, but once it does, it can be difficult to measure in the short term. It can be difficult to identify as a project on its own and more often than not, it is hiding behind a slew of other efforts and use cases that are more obtainable.

Over time, a pattern is created where each new initiative adds yet another layer to the foundation that has been only partially addressed.

This is especially true for the widespread utilization of artificial intelligence. The foundation that has new levels added to it for these initiatives becomes the basis for each new initiative aimed at addressing the norm of the organization.

This can often be applied incrementally, even starting with these low-hanging fruit domains that provide the most value.

Over time, this builds a level of consistency that can support use across a wider functional breadth.

Out of experience gained from the need to keep operations running, all organizations at this scale develop a version of this. It emerges as a result of the need to keep work moving while the level of constraint in which users are working is constantly changing.

A pilot demonstrates that a use case operates within a specific set of conditions, and the potential for scaling hinges on whether those conditions can be maintained as the model transitions into progressively more independently developed environments, which is where the actual work lies.

Context and Sources

JAMA Network Open and Nature Medicine have published research on model performance discrepancy across institutions. The NIST AI Risk Management Framework addresses data quality, representativeness, and context as core risk factors.

Christopher Hutchins
Founder & CEO, Hutchins Data Strategy Consultants

Recommended for you