Validity and Reliability of Two Abbreviated Versions of the Gross Motor Function Measure
Bibliographic record
Abstract
BACKGROUND: The "gold standard" for measuring gross motor function in children with cerebral palsy is the 66-item Gross Motor Function Measure (GMFM-66). OBJECTIVE: The purpose of this study was to estimate the validity and reliability of 2 abbreviated versions of the GMFM-66; one version involves an item set approach, and the other version involves a basal and ceiling approach. DESIGN: This was a measurement study comprising concurrent validity, comparability, and test-retest reliability components. METHODS: The study participants were 26 children who were 2 to 6 years of age and had cerebral palsy across all Gross Motor Function Classification System levels. In the first session, both abbreviated versions were administered by 2 independent raters; next, the full GMFM-66 was administered. In the second session, only the abbreviated versions were administered by the same raters. Concurrent validity, comparability of versions, and test-retest reliability were determined with intraclass correlation coefficients [ICC (2,1)]. RESULTS: Both versions demonstrated high levels of validity, with an ICC of .99 (95% confidence interval=0.972-0.997), reflecting associations with the GMFM-66. Both versions also were shown to be highly reliable, with ICCs of greater than .98 (95% confidence interval=0.965-0.994). LIMITATIONS: A smaller-than-expected sample was recruited for this study and may be a potential limitation of the study. CONCLUSION: Both versions of the GMFM-66 can be used in clinical practice or research. However, the GMFM-66 with the basal and ceiling approach is recommended as the preferred abbreviated version.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".