Evaluation of the revised Nipissing District Developmental Screening (NDDS) tool for use in general population samples of infants and children
Bibliographic record
Abstract
BACKGROUND: There is widespread interest in identification of developmental delay in the first six years of life. This requires, however, a reliable and valid measure for screening. In Ontario, the 18-month enhanced well-baby visit includes province-wide administration of a parent-reported survey, the Nipissing District Developmental Screening (NDDS) tool, to facilitate early identification of delay. Yet, at present the psychometric properties of the NDDS are largely unknown. METHOD: 812 children and their families were recruited from the community. Parents (most often mothers) completed the NDDS. A sub-sample (n = 111) of parents completed the NDDS again within a two-week period to assess test-retest reliability. For children 3 or younger, the criterion measure was the Bayley Scales of Infant Development, 3rd edition; for older children, a battery of other measures was used. All criterion measures were administered by trained assessors. Mild and severe delays were identified based on both published cut-points and on the distribution of raw scores. Sensitivity, specificity, positive and negative predictive values were calculated to assess agreement between tests. RESULTS: Test-retest reliability was modest (Spearman's rho = .62, p < 001). Regardless of the age of the child, the definition of delay (mild versus severe), or the cut-point used on the NDDS, sensitivities (from 29 to 68 %) and specificities (from 58 to 88 %) were poor to moderate. CONCLUSION: The modest test-retest results, coupled with the generally poor observed agreement with criterion measures, suggests the NDDS should not be used on its own for identification of developmental delay in community or population-based settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".