Somatic mosaicism detected by genome-wide sequencing in 500 parent–child trios with suspected genetic disease: clinical and genetic counseling implications
Bibliographic record
Abstract
Identifying genetic mosaicism is important in establishing a diagnosis, assessing recurrence risk, and providing accurate genetic counseling. Next-generation sequencing has allowed for the identification of mosaicism at levels below those detectable by conventional Sanger sequencing or chromosomal microarray analysis. The CAUSES Clinic was a pediatric translational trio-based genome-wide (exome or genome) sequencing study of 500 families (531 children) with suspected genetic disease at BC Children's and Women's Hospitals. Here we present 12 cases of apparent mosaicism identified in the CAUSES cohort: nine cases of parental mosaicism for a disease-causing variant found in a child and three cases of mosaicism in the proband for a de novo variant. In six of these cases, there was no evidence of mosaicism on Sanger sequencing-the variant was not detected on Sanger sequencing in three cases, and it appeared to be heterozygous in three others. These cases are examples of six clinical manifestations of mosaicism: a proband with classical clinical features of mosaicism (e.g., segmental abnormalities of skin pigmentation or asymmetrical growth of bilateral body parts), a proband with unusually mild manifestations of a disease, a mosaic proband who is clinically indistinguishable from the constitutive phenotype, a mosaic parent with no clinical features of the disease, a mosaic parent with mild manifestations of the disease, and a family in which both parents are unaffected and two siblings have the same disease-causing constitutional mutation. Our data demonstrate the importance of considering the possibility of mosaicism whenever exome or genome sequencing is performed and that its detection via genome-wide sequencing can permit more accurate genetic counseling.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".