Genomic and Developmental Models to Predict Cognitive and Adaptive Outcomes in Autistic Children
Bibliographic record
Abstract
Importance: Although early signs of autism are often observed between 18 and 36 months of age, there is considerable uncertainty regarding future development. Clinicians lack predictive tools to identify those who will later be diagnosed with co-occurring intellectual disability (ID). Objective: To predict ID in children diagnosed with autism. Design, Setting, and Participants: This prognostic study involved the development and validation of models integrating genetic variants and developmental milestones to predict ID. Models were trained, cross-validated, and tested for generalizability across 3 autism cohorts: Simons Foundation Powering Autism Research (SPARK), Simons Simplex Collection, and MSSNG. Autistic participants were assessed older than 6 years of age for ID. Study data were analyzed from January 2023 to July 2024. Exposures: Ages at attaining early developmental milestones, occurrence of language regression, polygenic scores for cognitive ability and autism, rare copy number variants, de novo loss-of-function and missense variants impacting constrained genes. Main Outcomes and Measures: The out-of-sample performance of predictive models was assessed using the area under the receiver operating characteristic curve (AUROC), positive predictive values (PPVs), and negative predictive values (NPVs). Results: A total of 5633 autistic participants (4574 male [81.2%]) were included in this analysis. On average, participants were diagnosed with autism at 4 (IQR, 3-7) years of age and assessed for ID at 11 (8-14) years of age, with 1159 participants (20.6%) being diagnosed with ID. The model integrating all predictors yielded an AUROC of 0.653 (95% CI, 0.625-0.681), and this predictive performance was cross-validated and generalized across cohorts. This modest performance reflected that only a subset of individuals carried large-effect variants, high polygenic scores, or presented delayed milestones. However, combinations of genetic variants that are typically not considered clinically relevant by diagnostic laboratories achieved PPVs of 55% and correctly identified 10% of individuals developing ID. The addition of polygenic scores to developmental milestones specifically improved NPVs rather than PPVs. Notably, the ability to stratify ID probabilities using genetic variants was up to 2-fold higher in individuals with delayed milestones compared with those with typical development. Conclusions and Relevance: Results of this prognostic study suggest that the growing number of neurodevelopmental condition-associated variants cannot, in most cases, be used alone for predicting ID. However, models combining different classes of variants with developmental milestones provide clinically relevant individual-level predictions that could be useful for targeting early interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".