Establishing Bayley-III cut-off scores at 21 months for predicting low IQ scores at 3 years of age in a preterm cohort
Bibliographic record
Abstract
OBJECTIVE: To evaluate predictive validity and establish cut-off scores on the Bayley-III at age 21 months that best predict Intelligence Quotient (IQ) scores <70 or <80) at 3 years in a high-risk preterm cohort. METHOD: Bayley-III evaluations at 21 months corrected age and intellectual assessments, primarily with the WPPSI-III, at 3 years corrected age were conducted with 520 infants born less than 29 weeks gestational age or less than 1250 g birth weight. Receiver Operator Characteristic (ROC) curves were used to establish Bayley-III Cognitive Composite cut-off scores that maximized Sensitivity and Specificity in predicting low IQ. Similar analyses were performed using the Language Composite, and a research derived mean Cognitive-Language Composite. RESULTS: =0.36). The ROC area under the Curve was 0.90 for the Cognitive Composite predicting IQ<70. The cut-off score that maximized Sensitivity and Specificity for predicting 3-year IQ<70 was a Cognitive Composite of <80. The ROC Area under the Curve was 0.80 for Cognitive Composites predicting IQ<80 and a Cognitive Composite cut-off score of <90 maximized Sensitivity and Specificity. CONCLUSION: In this high-risk preterm cohort, there was a strong association between the Bayley-III Cognitive Composite at 21 months and IQ at 3 years. A Cognitive Composite cut-off score of <80 optimized classification of IQ<70 at 3 years, and a Cognitive Composite cut-off score of <90 optimized classification of IQ<80.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".