Although clinical randomized trials (RCTs) are considered the gold standard for clinical evidence, they are not always feasible for certain medical interventions due to high cost, long timelines, and extended follow-up periods. Under these circumstances, real-world data (RWD) becomes increasingly important, as it can provide evidence that is not captured in randomized controlled trials. The authors analyzed current data sources in England and how to access them, addressing an area that has not been reviewed in detail in previous research.

Real-world data have two main advantages. First, researchers can use existing past data to study how often people get a disease, how easily they can get medical care, and how well treatments work. Second, they can use these data over time in the future to compare different treatments or evaluate whether changes in clinical practice improve patient outcomes.

The authors categorized data sources across England into three main types: Trusted Research Environments (TREs), national data repositories (such as the English Cancer Registry), and international collaborative platforms (such as federated learning systems).

Trusted Research Environments are secure systems that store health data from the UK National Health Service (NHS). Approved researchers can use this data for research, but they cannot download or take it out of the secure system. These systems include different types of health information, such as patient records, medical imaging, and other clinical data.

National data repositories are databases that focus on specific types of data, such as oncology or other disease areas, usually at a national level. These repositories organize and curate data to support research within specific medical fields.

International collaborative platforms are large-scale systems that enable data analysis across multiple countries, often using approaches such as federated learning. These platforms allow researchers to collaborate internationally while maintaining data privacy, and in some cases, they can link or coordinate analyses across different national datasets.

The study also compared local databases and international/national data infrastructures, focusing on how data are collected, categorized into specific groups (e.g., oncology patients), and securely transferred between systems while maintaining patient privacy. Structured categorization improves research efficiency. For example, in the English Cancer Registry, researchers can directly apply for access to cancer-specific datasets rather than searching through general administrative databases.

International databases are particularly useful for large-scale or comparative studies that require diverse datasets across populations. However, the authors note that effective use of these systems requires clearly defined research questions in order to easily find the right data.

Regardless of whether the database is local or national, these infrastructures are rapidly developing and increasingly important for cancer research. However, several challenges remain in building and maintaining them. First, protecting patient privacy is a top priority. Second, data access is often easier for researchers who are closely connected to the data providers, compared to those working remotely or externally. Third, maintaining these databases requires significant financial investment, robust security systems, and skilled personnel to manage and curate the data. Fourth, governance and compliance requirements for access can also pose challenges for some time-sensitive end users.

Reference

Price G, Peek N, Eleftheriou I, et al. An overview of real-world data infrastructure for cancer research. Clin Oncol (R Coll Radiol). 2025;38:103545. doi:10.1016/j.clon.2024.03.011

Leave a comment