Process-specific somatic mutation distributions vary with three-dimensional genome structure
Bibliographic record
Abstract
Abstract Somatic mutations arise during the life history of a cell. Mutations occurring in cancer driver genes may ultimately lead to the development of clinically detectable disease. Nascent cancer lineages continue to acquire somatic mutations throughout the neoplastic process and during cancer evolution (Martincorena and Campbell, 2015). Extrinsic and endogenous mutagenic factors contribute to the accumulation of these somatic mutations (Zhang and Pellman, 2015). Understanding the underlying factors generating somatic mutations is crucial for developing potential preventive, therapeutic and clinical decisions. Earlier studies have revealed that DNA replication timing (Stamatoyannopoulos et al., 2009) and chromatin modifications (Schuster-Böckler and Lehner, 2012) are associated with variations in mutational density. What is unclear from these early studies, however, is whether all extrinsic and exogenous factors that drive somatic mutational processes share a similar relationship with chromatin state and structure. In order to understand the interplay between spatial genome organization and specific individual mutational processes, we report here a study of 3000 tumor-normal pair whole genome datasets from more than 40 different human cancer types. Our analyses revealed that different mutational processes lead to distinct somatic mutation distributions between chromatin folding domains. APOBEC- or MSI-related mutations are enriched in transcriptionally-active domains while mutations occurring due to tobacco-smoke, ultraviolet (UV) light exposure or a signature of unknown aetiology (signature 17) enrich predominantly in transcriptionally-inactive domains. Active mutational processes dictate the mutation distributions in cancer genomes, and we show that mutational distributions shift during cancer evolution upon mutational processes switch. Moreover, a dramatic instance of extreme chromatin structure in humans, that of the unique folding pattern of the inactive X-chromosome leads to distinct somatic mutation distribution on X chromosome in females compared to males in various cancer types. Overall, the interplay between three-dimensional genome organization and active mutational processes has a substantial influence on the large-scale mutation rate variations observed in human cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".