Data feasibility in Healthcare: Comparing fit-for-purpose and fit-for-training

Healthcare organizations are investing heavily in both real-world evidence (RWE) studies and artificial intelligence (AI). At first glance, these two fields seem to rely on the same raw material: large healthcare datasets drawn from electronic health records (EHRs), claims databases, registries, genomics, imaging, and patient-generated data. Because the source data often overlaps, it is tempting to assume that if a dataset is suitable for one purpose, it can easily be repurposed for the other. However, this assumption is often imprecise.

 

The same oncology or hospital dataset may be ideal for answering a regulatory research question but fail completely as a training source for machine learning. Conversely, a dataset that powers a high-performing clinical AI tool may not be appropriate for epidemiologic or regulatory-grade evidence generation. Understanding the distinction matters because organizations increasingly want to reuse existing data assets across both domains.

 

Two different questions, one shared data landscape

 

At the core, RWE and AI are solving different problems.

 

RWE studies use real-world data (RWD) to answer a predefined scientific question. The goal is usually to estimate an association, compare treatment patterns, characterize a patient population, or evaluate outcomes outside controlled clinical trials. In this context, the key question is whether the dataset is fit-for-purpose. That means whether it contains the right variables, in sufficient quality, to answer one specific research question credibly.

 

AI implementation, by contrast, uses data to train models that learn patterns and make predictions. The goal is not to answer a single question directly but to create a reusable predictive function that can be deployed operationally. Here, the relevant question is whether the dataset is fit-for-training.

 

The distinction may sound subtle, but it changes the entire assessment process.

 

What fit-for-purpose means in RWE

 

A fit-for-purpose assessment in RWE is driven by the study protocol. Investigators begin with a precise question and then determine whether a given dataset can support it.

 

For example, a pharmaceutical company may want to evaluate whether patients receiving one immunotherapy have longer overall survival than those receiving another in metastatic lung cancer. The assessment focuses on whether the data can identify the correct patient cohort, capture treatments accurately, record survival outcomes, and include enough clinical variables to adjust for confounding.

 

This makes RWE feasibility highly question-specific. A dataset can be suitable for one study and unsuitable for another, even within the same disease area. An oncology EHR database may be excellent for survival analyses but unusable for progression-free survival if progression events are not consistently documented.

 

The main concern is scientific validity. Can the data support a reliable estimate for the chosen endpoint?

 

What fit-for-training means in AI

 

AI feasibility starts from a different perspective. Instead of asking whether variables exist for a specific analysis, it asks whether the dataset contains enough usable signal to train a model that performs well on future unseen patients. This shifts attention to issues that may not matter much in traditional RWE.

 

A machine learning team will examine whether labels are sufficiently accurate, whether the data volume is large enough, whether rare outcomes are represented, and whether records are consistent across hospitals or time periods. The same dataset may have all the variables needed for a retrospective study but still fail because labels are too noisy or because the model would not generalize beyond the institution where the data was collected.

 

For AI, the question is not only whether the data describes past events. It is whether it enables statistical learning that remains reliable in real-world deployment.

 

Why the same dataset can pass one test and fail the other

 

This difference becomes clearer when looking at a practical example.

 

Consider an oncology dataset of structured oncology RWD derived from cancer clinics. For an RWE study, this dataset may be ideal for comparing treatment patterns between two therapies. It contains diagnosis information, treatment timelines, and mortality data, which may be sufficient to estimate survival outcomes and adjust for patient characteristics.

 

For an AI use case, imagine training a model to predict treatment response before therapy starts. The same dataset may not work. Imaging may be absent, genomic testing may only be available for a subset of patients, and outcome labels such as response progression may rely on manual abstraction from clinical notes. These limitations may severely impair model training.

 

The data source is the same, but the suitability changes because the intended use changes.

 

The overlap: where datasets can support both

 

Despite these differences, there is meaningful overlap. Some datasets can support both RWE and AI, especially when they have three characteristics: breadth, longitudinal structure, and high-quality outcome capture.

Large integrated datasets combine clinical records, pathology, molecular testing, and outcomes. Such multimodal data can be valuable for observational studies and for AI development.

A rich oncology dataset with longitudinal treatment histories may support:

  • Comparative effectiveness research for an RWE team
  • Risk prediction or treatment recommendation models for an AI team

 

The overlap is strongest when the dataset contains both structured variables needed for epidemiologic analyses and dense high-resolution signals needed for machine learning.

 

This is increasingly common in precision medicine, radiology, and digital pathology, where the same source can support both scientific studies and algorithm development.

 

The hidden challenge: different quality standards

 

Although the source may overlap, the quality standards are often different.

 

RWE teams are primarily concerned with bias. They ask whether variables are complete enough to reduce confounding, whether follow-up is sufficient, and whether endpoints are valid.

 

AI teams are primarily concerned with learnability. They ask whether patterns are stable, whether labels are trustworthy, and whether data collected in one hospital will behave similarly in another.

 

A variable that is acceptable for RWE may still be problematic for AI. For example, missing laboratory values can sometimes be addressed through statistical imputation in a study. In AI training, systematic missingness may become a learned signal itself and lead to misleading model behavior.

 

This is why reusing datasets across both domains requires more than simple repurposing. It requires separate evaluation frameworks.

 

Why organizations increasingly assess both together

 

Many healthcare organizations are now trying to maximize the value of their data assets. A hospital, pharmaceutical company, or health-tech startup may invest heavily in creating a longitudinal patient dataset and want to use it for multiple purposes.

This has led to a growing trend: dual feasibility assessment.

 

Instead of evaluating a dataset only for one use case, teams increasingly assess it across both dimensions:

  • Can it answer current research questions?
  • Can it support future AI applications?

 

This approach helps organizations prioritize data investments. If a dataset can support both regulatory-grade evidence generation and machine learning development, its strategic value is much higher.

Companies such as Owkin are built around this idea, combining federated learning and research collaborations to create data networks that support both clinical research and AI model development.

 

A practical way to think about the difference

 

The simplest distinction is this:

  • RWE asks whether a dataset can support a valid answer to one predefined question.
  • AI asks whether a dataset can support learning a function that can be reused repeatedly.

 

That means RWE is study-centric, while AI is deployment-centric. Table 1 shows a clear comparison.

 

 

The overlap lies in datasets that are rich enough to do both: large longitudinal sources with high-quality labels, broad clinical coverage, and standardized structure. These datasets are valuable not because they automatically serve both purposes, but because they can be evaluated and optimized for both.

 

The strategy is not choosing between RWE and AI. It is building data governance that recognize when the same source can drive both scientific evidence and intelligent systems.

By Nadia Barozzi

Passionate about data-driven insights and the advancement of Real World Evidence research, drug safety and pharmacovigilance.