Associations of tumour somatic mutations with cancer-associated venous thromboembolism
Bibliographic record
Abstract
Abstract Background Venous thromboembolism (VTE) is a common complication of cancer. Complex interactions between tumour biology and the haemostatic system may contribute to development of cancer-associated VTE. Objectives This study examined associations of somatic mutations with VTE in a large multi-cancer cohort. Methods We analysed paired tumour and germline whole genome sequence data and electronic health records from 12,507 cancer patients recruited to the Genomics England National Genomic Research Library, to evaluate associations of somatic mutations across 608 genes, overall tumour mutational burden (TMB) and 25 single base substitution (SBS) mutational signatures with VTE. Interactions between somatic mutations and a germline polygenic risk score for VTE were also assessed. Results In multivariable Cox regressions adjusted for age, sex and genetic ancestry, somatic mutations in four genes associated with higher rates of VTE at a false-discovery rate <0.1: CDKN2A (Hazard ratio, HR=1.62 [95% confidence interval, 1.23-2.13]) , KRAS (HR=1.31 [1.12-1.53]), PCDH15 (HR=1.48 [1.24-1.76]) and TP53 (HR=1.55 [1.38-1.73] ). SBS8, a common mutation signature of unknown aetiology, was also associated with higher rates of VTE (HR=1.39 [1.16-1.66]). In contrast, TMB ≥20 mutations/Mb, two DNA mismatch repair signatures (SBS6 and SBS26) and one rare signature of unknown aetiology (SBS19) associated with lower rates of VTE. Evidence for these associations remained robust after additional adjustment for tumour type, stage, and systemic anti-cancer treatment. Conclusions These findings support the hypothesis that tumour somatic mutations influence risk of VTE. This may provide insights into the pathophysiology of cancer-associated VTE and inform future efforts to improve clinical risk prediction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".