Bibliographic record
Abstract
SIR–We thank Dr Williams for reading our article so thoroughly, including the online supplementary materials, and for sharing her concerns about our study that compared the predictive validity of the Harris Infant Neuromotor Test (HINT) and the Alberta Infant Motor Scale (AIMS).1 We also thank the editors for this opportunity to report additional data to assist readers in their review of the predictive validity of these two tests. To clarify one of the concerns mentioned, the sample tested initially on the HINT and AIMS comprised 144 infants; at the 2-year assessment, 114 of those infants were assessed on the Bayley-II Motor Scale as the outcome measure. Dr Williams was correct in pointing out the number of participants reported in Tables III and IV should have been 114 rather than 144; we apologize for that error. As shown in the descriptors for Tables III and IV in the article, the respective categorical outcomes were mild delay (−1SD) and significant delay (−2SD) on the Bayley-II Motor Scale, categorical definitions provided in the Bayley-II manual.2 At 2 years of age (time 3), 55 infants showed mild delays and two had significant delays. With regard to outcome status at 3 years, as stated within the data analysis section: ‘Owing to participant attrition between times 3 and 4, only the BSID-II motor scale outcomes at time 3 were used as the criterion variable for the categorical analyses.’ The racial/ethnic identities of the non-Caucasian participants in this study were 0.7% Black, 7.6% Asian, 2.8% Native/First Nations, 5.6% East Indian, 2.8% mixed Caucasian, and 0.7% mixed non-Caucasian. Although we agree with Dr Williams that these proportions are probably not comparable to US demographics, a recent study by McCoy et al.3 comparing US and Canadian HINT normative data and examining differences between US white and non-white groups concluded that: ‘There were no significant differences between HINT total scores for US and Canadian infants or for US racial or ethnic groups.’ This statement lends further credence to our conclusion that the results of the predictive validity study1 may be transferable to US infants. Dr Williams admirably cited the importance of a selection bias within our study. Although we did not identify specifically that attrition represents a type of selection bias, we discussed the attrition in our sample at some length within the limitations outlined in the discussion section (see p. 466). As she requested, the proportions of important demographic variables for infants assessed and not assessed at time 3 are compared to baseline (see Table I). To assess potential bias, it is relevant to note greater attrition from the high-risk group and of male infants. The rate of attrition was reported on page 464; more specifically, the number of typical infants at baseline, times 2, 3, and 4 respectively, was 58, 54, 49, and 34, and the number of at-risk infants was 86, 77, 65, and 38. Again, we thank Dr Williams and the journal’s editor for enabling us to expand upon the information provided in the published study. We hope that readers will find the additional information helpful in assessing the predictive validity of the HINT and AIMS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.061 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.027 | 0.028 |
| Insufficient payload (model declined to judge) | 0.022 | 0.016 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".