Assessing the accuracy of the Modified Checklist for Autism in Toddlers: a systematic review and meta‐analysis
Bibliographic record
Abstract
AIM: The Modified Checklist for Autism in Toddlers (M-CHAT) could be appropriate for universal screening for autism spectrum disorder (ASD) at 18 months and 24 months. Validation studies, however, reported differences in psychometric properties across sample populations. This meta-analysis summarized its accuracy measures and quantified their change in relation to patient and study characteristics. METHOD: Four electronic databases (MEDLINE, PsycINFO, CINAHL, and Embase) were searched to identify articles published between January 2001 and May 2016. Bayesian regression models pooled study-specific measures. Meta-regressions covariates were age at screening, study design, and proportion of males. RESULTS: On the basis of the 13 studies included, the pooled sensitivity was 0.83 (95% credible interval [CI] 0.75-0.90), specificity was 0.51 (95% CI 0.41-0.61), and positive predictive value was 0.53 (95% CI 0.43-0.63) in high-risk children and 0.06 (95% CI <0.01-0.14) in low-risk children. Sensitivity was higher for screening at 30 months compared with 24 months. INTERPRETATION: Findings indicate that the M-CHAT performs with low to moderate accuracy in identifying ASD among children with developmental concerns, but there was a lack of evidence on its performance in low-risk children or at age 18 months. Clinicians should account for a child's age and presence of developmental concern when interpreting their M-CHAT score. WHAT THIS PAPER ADDS: The Modified Checklist for Autism in Toddlers (M-CHAT) performs with low-to-moderate accuracy in children with developmental concerns. There is limited evidence supporting its use at 18 months or in low-risk children.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.006 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".