Content validity of mobility measures in arthrogryposis multiplex congenita: engaging clinicians and people with lived experience
Bibliographic record
Abstract
Introduction Lower-extremity impairment is prevalent in children with Arthrogryposis multiplex congenita (AMC), frequently leading to mobility limitations. Without AMC-specific assessment tools, clinicians and researchers often employ tools that have not been formally validated for the AMC population. This study aims to establish the content validity of commonly used mobility measures in children with AMC following the COnsensus-based Standards for health Measurement INstruments (COSMIN) and the International Classification of Functioning, Disability, and Health (ICF) framework. Methods Items from the measures “Functional Mobility Scale (FMS), Gillette Functional Assessment Questionnaire (FAQ), Functional Independence Measure for Children (WeeFIM), and Patient-Reported Outcomes Measurement Information System (PROMIS)” were linked to the ICF categories using the refined linking rules of the ICF. Three raters conducted independent linking, and inter-rater reliability was calculated using the Kappa coefficient. An expert panel consisting of people with lived experience, clinicians and researchers reviewed the ICF codes identified by the raters and evaluated the comprehensibility, relevance, and comprehensiveness of the four measures using the COSMIN standards. The Content Validity Index (CVI) and modified Kappa ( k *) were calculated. Results Inter-rater agreement was substantial [ κ = 0.79, (95% CI: 0.78–0.84)]. Most concepts (84.4%) were linked to the “Activities and Participation” domain, with a limited representation of “Environmental Factors” (8.9%) and “Body Functions” (6.7%). The CVI and k * values for most measures indicated excellent content validity (0.91 to 1), except for the PROMIS Mobility Young Adult (≤0.82). The expert panel found that measures exhibited high comprehensibility and relevance, but comprehensiveness was insufficient. Most studied mobility measures missed concepts such as pain, fatigue, mobility aids, and compensatory strategies. Conclusions FMS, FAQ, WeeFIM, and PROMIS (Parent Proxy/Pediatric) demonstrated good content validity. However, none of these measures fully address the full spectrum of mobility experiences in children with AMC. Incorporating missing concepts, such as environmental challenges, compensatory strategies, and pain, into existing or newly developed assessment tools is essential for providing a more holistic evaluation of functional mobility. Doing so will support more comprehensive clinical assessment, improve outcome tracking, and enhance care for children living with AMC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".