A Developmental Surveillance Score for Quantitative Monitoring of Early Childhood Milestone Attainment: Algorithm Development and Validation
Bibliographic record
Abstract
BACKGROUND: Developmental surveillance, conducted routinely worldwide, is fundamental for timely identification of children at risk of developmental delays. It is typically executed by assessing age-appropriate milestone attainment and applying clinical judgment during health supervision visits. Unlike developmental screening and evaluation tools, surveillance typically lacks standardized quantitative measures, and consequently, its interpretation is often qualitative and subjective. OBJECTIVE: Herein, we suggested a novel method for aggregating developmental surveillance assessments into a single score that coherently depicts and monitors child development. We described the procedure for calculating the score and demonstrated its ability to effectively capture known population-level associations. Additionally, we showed that the score can be used to describe longitudinal patterns of development that may facilitate tracking and classifying developmental trajectories of children. METHODS: We described the Developmental Surveillance Score (DSS), a simple-to-use tool that quantifies the age-dependent severity level of a failure at attaining developmental milestones based on the recently introduced Israeli developmental surveillance program. We evaluated the DSS using a nationwide cohort of >1 million Israeli children from birth to 36 months of age, assessed between July 1, 2014, and September 1, 2021. We measured the score's ability to capture known associations between developmental delays and characteristics of the mother and child. Additionally, we computed series of the DSS in consecutive visits to describe a child's longitudinal development and applied cluster analysis to identify distinct patterns of these developmental trajectories. RESULTS: The analyzed cohort included 1,130,005 children. The evaluation of the DSS on subpopulations of the cohort, stratified by known risk factors of developmental delays, revealed expected relations between developmental delay and characteristics of the child and mother, including demographics and obstetrics-related variables. On average, the score was worse for preterm children compared to full-term children and for male children compared to female children, and it was correspondingly worse for lower levels of maternal education. The trajectories of scores in 6 consecutive visits were available for 294,000 children. The clustering of these trajectories revealed 3 main types of developmental patterns that are consistent with clinical experience: children who successfully attain milestones, children who initially tend to fail but improve over time, and children whose failures tend to increase over time. CONCLUSIONS: The suggested score is straightforward to compute in its basic form and can be easily implemented as a web-based tool in its more elaborate form. It highlights known and novel relations between developmental delay and characteristics of the mother and child, demonstrating its potential usefulness for surveillance and research. Additionally, it can monitor the developmental trajectory of a child and characterize it. Future work is needed to calibrate the score vis-a-vis other screening tools, validate it worldwide, and integrate it into the clinical workflow of developmental surveillance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.034 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".