Rare Disease Registries: How Much Data Is Enough? Balancing Scientific Rigor with Clinical Reality

Rare diseases create a unique challenge for evidence generation. The number of patients is often small, diseases are heterogeneous, and clinical knowledge can be limited. Traditional clinical trials may struggle to recruit enough participants, follow patients for a sufficient period of time, or capture the complexity of disease progression in real life.

 

For these reasons, disease registries have become a key part of the rare disease evidence research. They provide an opportunity to collect longitudinal data from patients receiving care in different healthcare settings and can support many important activities, from understanding disease natural history to evaluating treatment effectiveness and informing regulatory and reimbursement decisions.

 

However, as expectations for real-world evidence (RWE) increase, registries are becoming increasingly complex. More variables, more domains, and more detailed assessments are often considered necessary to generate high-quality evidence or to satisfy the needs of a particular study protocol.

 

This creates an important question: how much data is enough? This is not so easy to answer, and the answer requires balancing two different perspectives: scientific rigor and clinical reality.

 

 

The case for standardized minimum datasets

 

A well-designed registry requires a clear understanding of what information should be collected and why.

 

Rare disease registries have historically suffered from fragmented and inconsistent data collection. Different centers may record different clinical outcomes, use different definitions, or follow different assessment approaches. When data are collected in different ways, combining information across countries and healthcare systems becomes challenging.

 

For this reason, many registry initiatives start with the development of a minimum dataset. These efforts usually involve clinicians, researchers, patients, methodologists, and other stakeholders who define the most important information needed to answer future research questions.

 

A recent example is the development of a minimum dataset for an international registry in pyoderma gangrenosum, a rare inflammatory skin disease. Through a Delphi consensus process, an international group of stakeholders reviewed 143 potential items and agreed on the inclusion of 118 items across 26 domains (1).

 

This type of work has significant value. A standardized dataset can improve consistency, facilitate collaboration between centers, support comparative analyses, and create stronger evidence for diseases. However, defining what should be collected is only one part of the challenge. The next question is whether healthcare professionals can realistically collect all this information during routine care.

 

 

The operational reality behind the registry spreadsheet

 

A registry exists because patients and healthcare professionals contribute data. The quality of the registry depends on their ability to provide accurate and complete information over time. However, in clinical practise, it has to be considered that every additional variable added to a registry has consequences. More fields require more time during clinical visits, more documentation, more training, and more resources.

 

In rare diseases, this challenge can be particularly significant. Many patients are managed in specialized centers where clinical teams already have limited time and resources. Adding extensive data collection requirements may compete with clinical priorities. There is also a difference between a variable being scientifically valuable and being consistently measurable in routine practice.

 

For example, a specific clinical assessment may be highly relevant for understanding disease progression. However, if it requires specialized training, additional equipment, or significant time during every visit, its completion may vary between centers. A registry with many variables collected inconsistently may provide less reliable evidence than a simpler registry with fewer variables collected with high completeness.

 

 

When missing data become systematic bias

 

Missing data are one of the biggest challenges in real-world data (RWD). In registries, missing information is often influenced by how healthcare is delivered, who provides care, and which patients are followed and this data missingness can create systematic bias.

 

For example, patients with more severe disease may have more frequent healthcare interactions and therefore more complete documentation. Patients with milder disease may have fewer visits and fewer recorded outcomes. Over time, the registry may create a picture of disease progression that reflects the patients who are most closely monitored.

 

International registries face additional challenges. Healthcare systems differ in access to diagnostic tests, specialist assessments, and clinical resources. A variable that is routinely collected in one country may rarely be available in another. This creates an important consideration: a mandatory dataset may unintentionally limit participation from certain regions or healthcare settings.

 

The most complete data may come from highly specialized academic centers with dedicated resources. These centers provide valuable expertise, but their patient populations may differ from the broader real-world population.

 

 

Minimum dataset versus minimum viable dataset

 

A useful concept for future registry design could be the distinction between a minimum dataset and a minimum viable dataset.

 

A minimum dataset defines the information needed from a scientific perspective. It answers the question: what data do we need to answer the research question?

A minimum viable dataset adds another important consideration: what data can realistically be collected with sufficient quality across different healthcare settings?

 

Both perspectives are needed. A registry should aim for scientific completeness, while also considering feasibility, sustainability, and long-term participation. This could mean moving towards more flexible registry models.

 

A core dataset could capture essential information for all patients. Additional modules could then be added depending on specific research objectives, treatments, disease stages, or emerging questions. This approach could reduce unnecessary burden while maintaining the ability to generate valuable evidence.

 

 

Towards smarter registry design

 

Rare disease registries might need better integration of different data sources and more efficient approaches to data collection. Electronic health records (EHRs), claims data, laboratory systems, patient-reported outcomes, and digital health technologies can help reduce manual data entry and provide additional information.

 

Artificial intelligence (AI) and natural language processing (NLP) may also support the extraction of relevant clinical information from medical records, although careful validation will remain essential.

 

Another important aspect is designing registries around clear evidence needs. Every variable included in a registry should have a purpose. Does it support understanding disease progression? Does it help evaluate treatment effectiveness? Does it address a regulatory or payer question? Does it improve patient care?

 

A registry becomes stronger when every piece of collected information contributes to a defined objective.

 

 

The future challenge: collecting the right data from enough patients

 

Rare disease registries represent an essential investment in improving evidence generation. They create opportunities that would otherwise be difficult or impossible because of small patient populations and limited clinical trial data. At the same time, registry success depends on a careful balance. Scientific ambition needs to be matched with clinical feasibility. The goal should be creating datasets that are meaningful, sustainable, and representative of the patients they aim to support.

 

The future of rare disease registries might depend less on collecting the largest possible amount of information and more on collecting the right information consistently over time.

 

High-quality evidence starts with high-quality data generation. The challenge is designing systems that allow patients, clinicians, and researchers to achieve this together.

 

 

Five questions to ask before adding a variable to a rare disease registry

 

Designing a registry requires continuous decisions about what information should be collected. Each additional variable has a cost, both for the healthcare professionals entering the data and for the patients whose care generates the information. Before adding a new field, registry designers could consider five simple questions.  

 

1. What decision will this data support? Every variable should have a clear purpose. The value of a data point depends on how it contributes to answering an important question. Will this information help understand disease progression? Evaluate treatment effectiveness? Support regulatory or reimbursement decisions? Improve clinical management? A variable without a clear future use may increase burden without improving the evidence generated.

 

2. Can this information be collected consistently in routine care? A variable may be scientifically relevant, but its usefulness depends on whether it can be measured reliably across different healthcare settings. Before including a field, it is important to understand how it will be collected in specialized centers, community settings, and different countries. Consistency often has greater value than complexity.

 

3. What is the expected data completeness? A registry variable is valuable only when enough patients have reliable information available. Low completion rates can limit analytical possibilities and may introduce bias. Understanding the expected level of completeness should be part of the decision to include a variable.   The question should not only be: "Can we collect this data?" but also: "Can we collect this data well over time?"

 

4. Could this information be captured through another source? The increasing availability of electronic health records, claims databases, laboratory systems, patient-reported outcomes, and digital technologies creates opportunities to reduce manual data collection. Before adding a new field for direct entry, it is worth considering whether the information already exists elsewhere and whether it can be integrated through appropriate methods.  

 

5. Does the value justify the burden? Every additional field represents a trade-off. More information can improve scientific understanding, but excessive complexity can reduce participation, increase missingness, and affect long-term sustainability. The strongest registries will be those that find the right balance between depth of information and feasibility of implementation. Rare disease evidence generation requires ambitious thinking, but also practical solutions. The best registry is not necessarily the one that collects the most data. It is the one that collects meaningful data consistently, from enough patients, over a long period of time.  

 

Reference:

  1. Haddadin OM, Jacobson ME, Becker SL, Chen D, Croitoru DO, Dissemond J, Renato V Gontijo J, Hampton PJ, Kelly RI, Marzano AV, Tada Y, Gerbens LAA, Ortega-Loayza AG. Minimum dataset for treatment effectiveness in pyoderma gangrenosum for an international registry: an international multidisciplinary eDelphi consensus. Br J Dermatol. 2026 Jul 17;195(2):270-279. doi: 10.1093/bjd/ljag084. PMID: 41784109.

 

By Nadia Barozzi

Passionate about data-driven insights and the advancement of Real World Evidence research, drug safety and pharmacovigilance.