Clinical potential of whole-genome data linked to mortality statistics in patients with breast cancer in the UK: a retrospective analysis
Bibliographic record
Abstract
BACKGROUND: Breast cancer is the most frequently diagnosed cancer in women. Survival is generally considered favourable, yet some patients remain at risk of early death. We aimed to assess whether comprehensive whole-genome sequencing (WGS) linked to mortality data could add prognostic value to existing clinical measures and identify patients who might respond to targeted therapeutics. METHODS: In this integrative, retrospective analysis, we analysed 2445 breast cancer tumours (any stage and molecular subtype) collected from 2403 patients recruited through 13 National Health Service Genomic Medicine Centres or hospitals in England affiliated to the 100 000 Genomes Project (100kGP) between 2012 and 2018. We linked 2208 (90%) cases with clinical data; mortality data were obtained for 1188 patients. Following high-depth WGS of tumour and matched normal DNA, we performed comprehensive WGS profiling seeking driver mutations, mutational signatures, and compound algorithmic scores for homologous recombination repair deficiency (HRD), mismatch repair deficiency, and tumour mutational burden. Data from 1803 additional patients with breast cancer from three independent cohorts were used to validate various findings. To evaluate the prognostic value of WGS features, we performed univariable and multivariable Cox regression on data from patients with stage I-III, ER-positive, HER2-negative breast cancer with a cancer-specific mortality endpoint (around 5-year follow-up). FINDINGS: Among 2445 tumours in the 100kGP breast cancer cohort, we observed genomic characteristics with immediate personalised medicine potential in 656 (26·8%), including features reporting HRD (298 [12·2%] total cases and 76 [6·3%] ER-positive, HER2-negative cases), highly individualised driver events, mutations underpinning resistance to endocrine therapy, and mutational signatures indicating therapeutic vulnerabilities. 373 (15·2%) cases had WGS features with potential for translational research, including compromised base excision repair and non-homologous end-joining dependency. Structural variation burden (hazard ratio 3·9 [95 CI% 2·4-6·2]; p<0·0001), high levels of APOBEC signatures (2·5 [1·6-4·1]; p<0·0001), and TP53 drivers (3·9 [2·4-6·2]; p<0·0001) were independently prognostic of customary clinical measures (age at diagnosis, stage, and grade) in patients with ER-positive, HER2-negative breast cancer. We developed a prognosticator for ER-positive, HER2-negative breast cancer capable of identifying patients who require either increased intervention or therapy de-escalation, validating the framework in the independent Swedish Cancerome Analysis Network-Breast (SCAN-B) dataset. INTERPRETATION: We show that breast cancer genomes are rich in predictive and prognostic value. We propose a two-step model for effective clinical application. First, the identification of candidates for targeted therapies or clinical trials using highly individualised genomic markers. Second, for patients without such features, the implementation of enhanced prognostication using genomic features alongside existing clinical decision-making factors. FUNDING: National Institute of Health Research, Breast Cancer Research Foundation, Dr Josef Steiner Cancer Research Award 2019, Basser Gray Prime Award 2020, Cancer Research UK, Sir Jeffrey Cheah Early Career Fellowship, the Mats Paulsson Foundation, the Fru Berta Kamprads Foundation, and the Swedish Research Council.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".