Exploring the genetic architecture of autism spectrum disorder
Bibliographic record
Abstract
Autism spectrum disorder (ASD) is one of the most common neurodevelopmental disorders affecting approximately 1% of the global population.Individuals with ASD have impairments in social interaction and communication, and restricted and repetitive patterns of behaviour, interests, and activities.The spectrum of clinical presentation can vary between subjects with ASD and most individuals have at least one co-occurring psychiatric or medical condition.The early onset (<3 years), lifelong persistence, economic burden, and lack of therapeutics for ASD warrants a need to elucidate its etiology and in turn, mitigate its impact.The role of genetic risk factors in ASD have been firmly established, with an estimated heritability of 65-91%, and an accompanying complex genetic architecture.Like other complex traits, different classes of genetic variants with varying frequencies and penetrance have been implicated in the genetic liability for ASD.Early efforts focused on identifying rare, large-effect genetic risk factors.Yet, there remains insufficient evidence for ASD-specific genes to date.Gaining insight into the complex genetic architecture of ASD is fundamental to understanding the mechanisms of the disorder.Today, large-scale genetic and clinical data are available to power statistical models of the disorder.Yet, the clinical and genetic heterogeneity of the disorder remains a challenge to overcome despite the substantial increases in sample size over the last decade.As such, approaches that leveragerather than reducethe complexity of ASD can be useful to uncover the genetic underpinnings of the disorder.In this thesis, we used the genetic and clinical data of tens of thousands of individuals from families with ASD and the general population to characterize the heterogeneous risk factors of ASD.First, we identified a rare inherited copy-number variant (CNV) encompassing the CNTN5 gene shared by all four affected brothers of a multiplex family with ASD.We validated the role of this variant in a large case-control study, which confirmed its association with ASD risk and other neuropsychiatric conditions.This study underscored the importance of characterizing variants of intermediate effect size in the etiology of ASD to elucidate its complete genetic architecture.Second, we explored the genetic liability for ASD conferred through common variants.In this study, we combined multiple polygenic risk scores (PRSs) for ASD-related traits to capture the CH.Statistical analyses and result interpretation were performed by ZS, VRB, ED, JPR, PA, and CEC.Data analysis and manuscript writing was performed by ZS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".