Reliability of family report for the Gross Motor Function Classification System
Bibliographic record
Abstract
The aim of this study was to determine the reliability of family reports for the Gross Motor Function Classification System (GMFCS), a condition-specific discriminative measure of severity of movement disability for children with cerebral palsy (CP). We conducted a cross-sectional survey using a short questionnaire with families of children with CP for whom we already had ratings of GMFCS level made by a health professional. We assessed the potentially confounding effect of whether the family had discussed the GMFCS with a professional. Two hundred and one questionnaires were posted to families of which 97 (48%) were completed and returned. Mean age of the children (53 males, 40 females) was 9 years 5 months (SD 1 year 1 month), range 6 to 11 years. Children of the families who responded encompassed the spectrum of types and distribution of impairment and severity of movement disability. The intraclass correlation coefficient (ICC) of agreement between professionals and families who had discussed their child's GMFCS level with a health professional (n=35) was 0.97 (95% confidence interval [CI] 0.96 to 0.98); for those who had not (n=52) the ICC was 0.92 (95% CI 0.91 to 0.93); and for the whole sample (n=93) the ICC was 0.94 (95% CI 0.90 to 0.96). Stability between ratings made by health professionals for children when they were in the 4 to 6 year age band of the GMFCS and ratings made by families for the same children when they were in the 6 to 12 year age band (n=35) was ICC=0.96 (95% CI 0.95 to 0.97). The excellent agreement demonstrated in this study suggests that family reports of the GMFCS made by using our questionnaire provide a reliable method for measuring gross motor function in children between 6 and 12 years old. This might be more efficient for observational studies of large populations, experimental research, or community health administration than direct observation, particularly when professional assessment is not feasible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".