The Origins and Consequences of Localized and Global Somatic Hypermutation
Bibliographic record
Abstract
Abstract Cancer is a disease of the genome, but the dramatic inter-patient variability in mutation number is poorly understood. Tumours of the same type can differ by orders of magnitude in their mutation rate. To understand potential drivers and consequences of the underlying heterogeneity in mutation rate across tumours, we evaluated both local and global measures of mutation density: both single-stranded and double-stranded DNA breaks in 2,460 tumours of 38 cancer types. We find that SCNAs in thousands of genes are associated with elevated rates of point-mutations, while similarly point-mutation patterns in dozens of genes are associated with specific patterns of DNA double-stranded breaks. These candidate drivers of mutation density are enriched for known cancer drivers, and preferentially occur early in tumour evolution, appearing clonally in all cells of a tumour. To supplement this understanding of global mutation density, we developed and validated a tool called SeqKat to identify localized “rainstorms” of point-mutations (kataegis). We show that rates of kataegis differ by four orders of magnitude across tumour types, with malignant lymphomas showing the highest. Tumours with TP53 mutations were 2.6-times more likely to harbour a kataegic event than those without, and 239 SCNAs were associated with elevated rates of kataegis, including loss of the tumour-suppressor CDKN2A . We identify novel subtypes of kataegic events not associated with aberrant APOBEC activity, and find that these are localized to specific cellular regions, enriched for MYC-target genes. Kataegic events were associated with patient survival in some, but not all tumour types, highlighting a combination of global and tumour-type specific effects. Taken together, we reveal a landscape of genes driving localized and tumour-specific hyper-mutation, and reveal novel mutational processes at play in specific tumour types.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".