Early gross motor development: Agreement between the AIMS and the BSID-III
Bibliographic record
Abstract
Background: Early gross motor development is a crucial indicator of overall neurodevelopment. In low- and middle-income countries, lack of accessible assessment tools poses challenges for healthcare professionals evaluating infant neurodevelopment. Objectives: To determine the agreement between the Alberta Infant Motor Scale (AIMS) and Bayley Scales of Infant Development-III (BSID-III) gross motor domain at 6 months and to evaluate the predictive validity of the AIMS at 6 months for identifying severe gross motor delays at 18 months. Method: This nested subgroup study assessed 112 full-term infants using both AIMS and BSID-III at 6 months and BSID-III at 18 months. Agreement between measures was determined using Bland-Altman plots, while predictive validity was evaluated using receiver operating characteristic (ROC) curves with various cut-off scores. Results: Bland-Altman analysis showed strong agreement between AIMS and BSID-III in the lower-performance range, with bias only in scores above 33. The traditional 10th percentile AIMS cut-off had low sensitivity (27.3%) but high specificity (98%) for predicting delays at 18 months. A modified 23rd percentile cut-off improved sensitivity to 63.6% while maintaining acceptable specificity (81.6%), with a 95.2% negative predictive value (NPV). Conclusion: The AIMS demonstrates strong agreement with BSID-III when identifying potential developmental delays. The proposed 23rd percentile cut-off offers a more balanced screening threshold for this population. Clinical Implications: The AIMS presents a viable alternative to the BSID-III for initial screening in resource-limited settings. The high NPV at the 23rd percentile cut-off makes it useful for ruling out developmental delays.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".