Crossing borders: the need for empirical evidence of real-world evidence transportability in oncology
Bibliographic record
Abstract
WHAT IS THIS ARTICLE ABOUT?: This article discusses the challenges of using non local real-world evidence (RWE) in health technology assessments (HTA) when local data are unavailable, insufficient, or inappropriate. HTA organizations often prefer data collected locally or regionally, but the lack of suitable data in many markets has increased interest in understanding data 'transportability' - whether data from one country or population can be used to predict outcomes in another. Established in 2024, the Flatiron Fostering Oncology RWE Use Cases and Methods (FORUM) research consortium is exploring when and how non-local RWE can be effectively applied, with initial work focused on lung cancer, breast cancer and multiple myeloma. WHAT DOES THE EVIDENCE SUGGEST SO FAR?: Initial studies suggest RWE from the US could predict outcomes in other countries with proper adjustment for population and treatment differences. Recent research in advanced non-small cell lung cancer demonstrated that adjusted US data provided comparable survival to real observed outcomes in Canada and the UK. This limited evidence base indicates that non-local RWE can help inform decision-making when local data is unavailable. WHAT STUDIES ARE NEEDED NEXT?: The FORUM consortium is expanding research to other cancer types and countries to better understand RWE transportability. Future studies will focus on comparing outcomes across diverse healthcare systems, identifying key variables for adjustment and developing guidelines for when and how non-local data can be used. These efforts aim to create a framework for the use of global RWE in oncology HTA decision-making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.190 | 0.090 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".