Development and validation of item sets to improve efficiency of administration of the 66‐item Gross Motor Function Measure in children with cerebral palsy
Bibliographic record
Abstract
AIM: To develop an algorithmic approach to identify item sets of the 66-item version of the Gross Motor Function Measure (GMFM-66) to be administered to individual children, and to examine the validity of the algorithm for obtaining a GMFM-66 score. METHOD: An algorithmic approach was used to identify item sets of the GMFM-66 (GMFM-66-IS) using data from 95 males and 79 females with cerebral palsy (CP; mean age 14y 7mo, SD 1y 8mo, range 12y 7mo to 17y 8mo). The GMFM-66-IS scores were then validated using combined data from three Dutch studies involving 134 males and 92 females with CP (mean age 7y, SD 4y 6mo, range 1y 4mo to 13y 8mo), representing all levels of the Gross Motor Function Classification System. RESULTS: The final algorithm contains three decision items from the GMFM-66 that determine which one of four item sets to administer. The GMFM-66-IS has excellent agreement with the full GMFM-66 both at a single assessment (intraclass correlation coefficient [ICC]=0.994, 95% confidence intervals [CI] 0.993-0.996) and across repeat assessments (ICC=0.92, 95% CI 0.89-0.95). INTERPRETATION: The GMFM-66-IS is a promising alternative to the full GMFM-66. Users should be consistent in their choice of measure (GMFM-66 or GMFM-66-IS) on repeat testing and clearly identify which method was used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".