Food for Thought: Machine Learning in Autism Spectrum Disorder Screening of Infants
Bibliographic record
Abstract
Diagnoses of autism spectrum disorders (ASD) are typically made after toddlerhood by examining behavioural patterns. Earlier identification of ASD enables earlier intervention and better outcomes. Machine learning provides a data-driven approach of diagnosing autism at an earlier age. This review aims to summarize recent studies and technologies utilizing machine learning based strategies to screen infants and children under the age of 18 months for ASD, and identify gaps that can be addressed in the future. We reviewed nine studies based on our search criteria, which includes primary studies and technologies conducted within the last 10 years that examine children with ASD or at high risk of ASD with a mean age of less than 18 months old. The studies must use machine learning analysis of behavioural features of ASD as major methodology. A total of nine studies were reviewed, of which the sensitivity ranges from 60.7% to 95.6%, the specificity ranges from 50% to 100%, and the accuracy ranges from 60.9% to 97.7%. Factors that contribute to the inconsistent findings include the varied presentation of ASD among patients and study design differences. Previous studies have shown moderate accuracy, sensitivity and specificity in the differentiation of ASD and non-ASD individuals under the age of 18 months. The application of machine learning and artificial intelligence in the screening of ASD in infants is still in its infancy, as observed by the granularity of data available for review. As such, much work needs to be done before the aforementioned technologies can be applied into clinical practice to facilitate early screening of ASD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".