Discovery and diagnostic interpretation of germline and mosaic variation in developmental conditions
Bibliographic record
Abstract
Genome sequencing has provided paradigm shifting access to variability across humans. Sequencing technologies have discovered variants that can influence specific phenotypes and risk for disease. Despite this transformative technology, clinically interpretable diagnostic variants remain unknown for most rare and common disease patients as a result of the challenges associated with systematically analyzing and interpreting all variation in each human genome. The work presented here studies two early developmental disorders that are routinely referred for clinical genetic testing, autism spectrum disorder and fetal structural anomalies. In this thesis, we demonstrate that short-read genome sequencing can capture all variant classes identified by current clinical approaches and quantify the novel diagnoses contributing to these disorders. While germline variants account for most genetic diagnoses, postzygotic mutations, variants with low allele fraction present in only a subset of cells in the body, can contribute to disease yet are not systematically analyzed. We present a dataset of postzygotic mutations from the largest number of ASD samples and the first from standard genome sequencing. We first describe an approach to leverage genome sequencing for the diagnosis of autism spectrum disorder and fetal structural anomalies to replace the current clinical standard-of-care tests: microarray, karyotype, and exome sequencing. This work demonstrates that genome sequencing identifies more diagnostic variants than any single test or combination of tests. It captures all currently ascertainable variants as well as new variants unique to this technology, a category expected to increase as interpretation of genome sequencing variants matures. Then, we identify postzygotic mutations in an expanded cohort of 28,349 autism spectrum disorder cases and family members, providing a resource to analyze the impact of this class of variation on individuals with autism spectrum disorder compared to their unaffected siblings. Together, this work further delineates how genome sequencing can be implemented immediately as a first-line diagnostic test for autism spectrum disorder and prenatal anomalies while providing rationale to further explore the contribution of postzygotic mutations to the genetic etiology of these early developmental disorders.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".