Stability of the Gross Motor Function Classification System, Manual Ability Classification System, and Communication Function Classification System
Bibliographic record
Abstract
AIM: To determine the stability of the Gross Motor Function Classification System (GMFCS), Manual Ability Classification System (MACS), and Communication Function Classification System (CFCS) over 1-year and 2-year intervals using a process for consensus classification between parents and therapists. METHOD: Participants were 664 children with cerebral palsy (CP), 18 months to 12 years of age, one of their parents, and 90 therapists. Consensus between parents and therapists on level of function was ≥92% for the GMFCS, MACS, and CFCS. A linearly weighted kappa coefficient of ≥0.75 was the criterion for stability. RESULTS: Kappa coefficients varied from 0.76 to 0.88 for the GMFCS, 0.59 to 0.73 for the MACS, and 0.57 to 0.77 for the CFCS. For children younger than 4 years of age, level of function did not change for 58.2% on the GMFCS, 30.3% on the MACS, and 39.3% on the CFCS. For children 4 years of age or older, level of function did not change for 72.3% on the GMFCS, 49.1% on the MACS, and 55% on the CFCS. INTERPRETATION: The findings support repeated classification of children over time. The kappa coefficients for the GMFCS are attributed to descriptions of levels for each age band. Consensus classification facilitates discussion between parents and professionals that has implications for shared decision-making. WHAT THIS PAPER ADDS: The findings support repeated classification of children over time. Stability was higher for the Gross Motor Function Classification System than the Manual Ability Classification System and Communication Function Classification System. The function of younger children was more likely to be reclassified. Percentage agreement between parents and therapists using consensus classification varied from 92% to 97%. The intraclass correlation coefficient overestimated stability compared with the weighted kappa coefficient.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".