Diagnosis-informed connectivity subtyping discovers subgroups of autism with reproducible symptom profiles
Bibliographic record
Abstract
Clinical heterogeneity has been one of the main barriers to develop effective biomarkers and therapeutic strategies in autism spectrum disorder (ASD). Recognizing this challenge, much effort has been made in recent neuroimaging studies to find biologically more homogeneous subgroups (called 'neurosubtypes') in autism. However, most approaches have rarely evaluated how much the employed features in subtyping represent the core anomalies of ASD, obscuring its utility in actual clinical diagnosis. To address this, we combined two data-driven methods, 'connectome-based gradient' and 'functional random forest', collectively allowing to discover reproducible neurosubtypes based on resting-state functional connectivity profiles that are specific to ASD. Indeed, the former technique provides the features (as input for subtyping) that effectively summarize whole-brain connectome variations in both normal and ASD conditions, while the latter leverages a supervised random forest algorithm to inform diagnostic labels to clustering, which makes neurosubtyping driven by the features of ASD core anomalies. Applying this framework to the open-sharing Autism Brain Imaging Data Exchange repository data (discovery, n = 103/108 for ASD/typically developing [TD]; replication, n = 44/42 for ASD/TD), we found three dominant subtypes of functional gradients in ASD and three subtypes in TD. The subtypes in ASD revealed distinct connectome profiles in multiple brain areas, which are associated with different Neurosynth-derived cognitive functions previously implicated in autism studies. Moreover, these subtypes showed different symptom severity, which degree co-varies with the extent of functional gradient changes observed across the groups. The subtypes in the discovery and replication datasets showed similar symptom profiles in social interaction and communication domains, confirming a largely reproducible brain-behavior relationship. Finally, the connectome gradients in ASD subtypes present both common and distinct patterns compared to those in TD, reflecting their potential overlap and divergence in terms of developmental mechanisms involved in the manifestation of large-scale functional networks. Our study demonstrated a potential of the diagnosis-informed subtyping approach in developing a clinically useful brain-based classification system for future ASD research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".