Somatic point mutations occurring early in development: a monozygotic twin study
Bibliographic record
Abstract
The identification of somatic driver mutations in cancer has enabled therapeutic advances by identifying drug targets critical to disease causation. However, such genomic discoveries in oncology have not translated into advances for non-cancerous disease since point mutations in a single cell would be unlikely to cause non-malignant disease. An exception to this would occur if the mutation happened early enough in development to be present in a large percentage of a tissue's cellular population. We sought to identify the existence of somatic mutations occurring early in human development by ascertaining base-pair mutations present in one of a pair of monozygotic twins, but absent from the other and assessing evidence for mosaicism. To do so, we genome-wide genotyped 66 apparently healthy monozygotic adult twins at 506 786 high-quality single nucleotide polymorphisms (SNPs) in white blood cells. Discrepant SNPs were verified by Sanger sequencing and a selected subset was tested for mosaicism by targeted high-depth next-generation sequencing (20 000-fold coverage) as a surrogate marker of timing of the mutation. Two de novo somatic mutations were unequivocally confirmed to be present in white blood cells, resulting in a frequency of 1.2×10(-7) mutations per nucleotide. There was little evidence of mosaicism on high-depth next-generation sequencing, suggesting that these mutations occurred early in embryonic development. These findings provide direct evidence that early somatic point mutations do occur and can lead to differences in genomes between otherwise identical twins, suggesting a considerable burden of somatic mutations among the trillions of mitoses that occur over the human lifespan.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".