High-precision machine learning identifies a reproducible functional connectivity signature of autism spectrum diagnosis in a subset of individuals
Bibliographic record
Abstract
BACKGROUND: Discovery of predictive biomarkers is essential for understanding the neurobiological underpinnings of autism spectrum diagnosis (ASD) and improving identification. Resting-state functional connectivity analyses of individuals with ASD have established sensitivity of brain connectivity at the group level. However, the extensive heterogeneity in ASD limits the translation of these findings into reliable individual-level biomarkers. We analyzed the Autism Brain Imaging Data Exchange 1 and 2 datasets, calculating Pearson's correlation-based functional connectivity across 18 brain networks. Using transductive conformal prediction, a machine learning approach that assigns confidence scores to predictions based on conformality to known classes, we classified individuals with ASD and neurotypical controls. RESULTS: By combining predictors into an ensemble using hierarchical agglomerative clustering, we identified a signature that confers a more than 7-fold increase in individual risk of ASD, yet is still identified in an estimated 1 in 200 individuals in the general population. The individual risk conferred by the model is increased 4-fold over that of previously published imaging models and outperforms the current state of the art in precision for ASD classification. The high-risk signature was characterized by underconnectivity of transmodal brain networks, including the frontoparietal and basal ganglia network, and subcomponents of the limbic and default mode networks. CONCLUSIONS: A highly targeted prediction model can identify a subset of functional connectivity alterations that confer high risk for ASD at the individual level, which may be masked by traditional machine learning models due to ASD heterogeneity. Results could help disentangle the multitude of etiological pathways and behavioral symptoms that challenge our understanding of ASD by focusing on highly penetrant connectivity signatures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.031 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".