The complexity of finding fit-for-purpose real-world data for oncology patients with rare NTRK gene fusions and a novel solution
Bibliographic record
Abstract
Background Multiple data sources suggest low frequencies of neurotrophic tyrosine receptor kinase ( NTRK ) gene fusions across common solid tumors, ranging from 0.18% to 0.30%, making it a relevant target for a tumor-agnostic development using a single-arm basket trial. The objective of this study was to explore a multifaceted approach to building a pooled real-world dataset of patients with select solid tumors harboring NTRK1-3 gene fusions from multiple clinicogenomic databases (CGDBs) and clinical site sources. Materials and methods A novel approach was explored for identifying patients with NTRK gene fusion-positive solid tumors from real-world data (RWD) sources through existing CGDBs and global clinical site surveys. CGDBs were assessed for inclusion based on patient data, genomic testing, and accessibility of patient-level data. The clinical site surveys were used to determine the practicality of designing a chart review study where patients were identified via electronic medical records and molecular assays. Results Approximately 19% of CGDBs and 1% of the clinical sites included in the survey outreach were eligible for inclusion. Combining data from CGDBs and clinical sites through a retrospective chart review yielded a real-world cohort of 512 patients with NTRK gene fusion-positive solid tumors. Conclusion The rarity of patients with NTRK gene fusion-positive cancer and number of eligible CGDBs/clinical sites present challenges for the identification of sufficient RWD sources that can be used in comparative effectiveness studies to contextualize results from single-arm trials. This approach provides a potential solution to identify fit-for-purpose RWD to support precision medicine for patients with cancer harboring rare genomic alterations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".