MétaCan
Menu
Back to cohort
Record W4415023189 · doi:10.1016/s1470-2045(25)00400-0

Clinical potential of whole-genome data linked to mortality statistics in patients with breast cancer in the UK: a retrospective analysis

2025· article· en· W4415023189 on OpenAlexaff
Daniella Black, Helen Davies, Gene Ching Chiek Koh, Lucia Chmelova, Marko Cubric, G. C. Chan, Andrea Degasperi, Jan Czarnecki, Ping Jing Toong, Yasin Memari, James Whitworth, Salome Jingchen Zhao, Yogesh Kumar, Shadi Basyuni, Giuseppe Rinaldi, Scott Shooter, Vladyslav Dembrovskyi, R. A. Davies, Maria Chatzou, Ellen Copson, Carlo Palmieri, Åke Borg, John C. Ambrose, Catey Bunce, Alona Sosinsky, Prabhu Arumugam, Matthew A. Brown, Johan Staaf, Nicholas C. Turner, Serena Nik‐Zainal

Bibliographic record

VenueThe Lancet Oncology · 2025
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsInstitute of Cancer Research
FundersCancer Research UK Cambridge Institute, University of CambridgeMedical Research CouncilFru Berta Kamprads StiftelseLunds UniversitetCancerfondenVetenskapsrådetSunway UniversityCancer Research UKNational Institute for Health and Care ResearchHeart of England NHS Foundation TrustDr. Josef Steiner KrebsstiftungDepartment of Health and Social CareMats Paulssons StiftelseWellcome TrustNIHR Cambridge Biomedical Research CentreBreast Cancer Research Foundation
KeywordsBreast cancerLinked dataCancerPopulationRetrospective cohort studyEpidemiology

Abstract

fetched live from OpenAlex

BACKGROUND: Breast cancer is the most frequently diagnosed cancer in women. Survival is generally considered favourable, yet some patients remain at risk of early death. We aimed to assess whether comprehensive whole-genome sequencing (WGS) linked to mortality data could add prognostic value to existing clinical measures and identify patients who might respond to targeted therapeutics. METHODS: In this integrative, retrospective analysis, we analysed 2445 breast cancer tumours (any stage and molecular subtype) collected from 2403 patients recruited through 13 National Health Service Genomic Medicine Centres or hospitals in England affiliated to the 100 000 Genomes Project (100kGP) between 2012 and 2018. We linked 2208 (90%) cases with clinical data; mortality data were obtained for 1188 patients. Following high-depth WGS of tumour and matched normal DNA, we performed comprehensive WGS profiling seeking driver mutations, mutational signatures, and compound algorithmic scores for homologous recombination repair deficiency (HRD), mismatch repair deficiency, and tumour mutational burden. Data from 1803 additional patients with breast cancer from three independent cohorts were used to validate various findings. To evaluate the prognostic value of WGS features, we performed univariable and multivariable Cox regression on data from patients with stage I-III, ER-positive, HER2-negative breast cancer with a cancer-specific mortality endpoint (around 5-year follow-up). FINDINGS: Among 2445 tumours in the 100kGP breast cancer cohort, we observed genomic characteristics with immediate personalised medicine potential in 656 (26·8%), including features reporting HRD (298 [12·2%] total cases and 76 [6·3%] ER-positive, HER2-negative cases), highly individualised driver events, mutations underpinning resistance to endocrine therapy, and mutational signatures indicating therapeutic vulnerabilities. 373 (15·2%) cases had WGS features with potential for translational research, including compromised base excision repair and non-homologous end-joining dependency. Structural variation burden (hazard ratio 3·9 [95 CI% 2·4-6·2]; p<0·0001), high levels of APOBEC signatures (2·5 [1·6-4·1]; p<0·0001), and TP53 drivers (3·9 [2·4-6·2]; p<0·0001) were independently prognostic of customary clinical measures (age at diagnosis, stage, and grade) in patients with ER-positive, HER2-negative breast cancer. We developed a prognosticator for ER-positive, HER2-negative breast cancer capable of identifying patients who require either increased intervention or therapy de-escalation, validating the framework in the independent Swedish Cancerome Analysis Network-Breast (SCAN-B) dataset. INTERPRETATION: We show that breast cancer genomes are rich in predictive and prognostic value. We propose a two-step model for effective clinical application. First, the identification of candidates for targeted therapies or clinical trials using highly individualised genomic markers. Second, for patients without such features, the implementation of enhanced prognostication using genomic features alongside existing clinical decision-making factors. FUNDING: National Institute of Health Research, Breast Cancer Research Foundation, Dr Josef Steiner Cancer Research Award 2019, Basser Gray Prime Award 2020, Cancer Research UK, Sir Jeffrey Cheah Early Career Fellowship, the Mats Paulsson Foundation, the Fru Berta Kamprads Foundation, and the Swedish Research Council.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.013
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.011
Threshold uncertainty score0.022

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.013
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.005
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.026
GPT teacher head0.367
Teacher spread0.342 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueThe Lancet OncologySame topicCancer Genomics and DiagnosticsFrench-language works237,207