Disease similarity network analysis of Autism Spectrum Disorder and comorbid brain disorders
Bibliographic record
Abstract
Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder with heterogeneous clinical presentation, variable severity, and multiple comorbidities. A complex underlying genetic architecture matches the clinical heterogeneity, and evidence indicates that several co-occurring brain disorders share a genetic component with ASD. In this study, we established a genetic similarity disease network approach to explore the shared genetics between ASD and frequent comorbid brain diseases (and subtypes), namely Intellectual Disability, Attention-Deficit/Hyperactivity Disorder, and Epilepsy, as well as other rarely co-occurring neuropsychiatric conditions in the Schizophrenia and Bipolar Disease spectrum. Using sets of disease-associated genes curated by the DisGeNET database, disease genetic similarity was estimated from the Jaccard coefficient between disease pairs, and the Leiden detection algorithm was used to identify network disease communities and define shared biological pathways. We identified a heterogeneous brain disease community that is genetically more similar to ASD, and that includes Epilepsy, Bipolar Disorder, Attention-Deficit/Hyperactivity Disorder combined type, and some disorders in the Schizophrenia Spectrum. To identify loss-of-function rare de novo variants within shared genes underlying the disease communities, we analyzed a large ASD whole-genome sequencing dataset, showing that ASD shares genes with multiple brain disorders from other, less genetically similar, communities. Some genes (e.g., SHANK3, ASH1L, SCN2A, CHD2 , and MECP2 ) were previously implicated in ASD and these disorders. This approach enabled further clarification of genetic sharing between ASD and brain disorders, with a finer granularity in disease classification and multi-level evidence from DisGeNET. Understanding genetic sharing across disorders has important implications for disease nosology, pathophysiology, and personalized treatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".