Translation, inter-rater reliability, agreement, and internal consistency of the Spanish version of the cumulated ambulation score in patients after hip fracture
Bibliographic record
Abstract
Purpose: To translate the Cumulated Ambulation Score into Spanish, and to examine its inter-rater reliability, agreement and internal consistency.Materials and Methods: Two occupational therapists independently used the Spanish version of the Cumulated Ambulation Score (three activities scored from 0–2 points) to assess 60 consecutive patients with hip fracture within the first post-surgery week at a traumatology service of a public hospital. We used linear weighted kappa (κ) statistics to determine inter-rater reliability, percent agreement to assess measurement error, Cronbach’s α coefficient to establish the internal consistency, and the McNemar–Bowker test to evaluate for systematic between-rater differences.Results: The κ was ≥ 0.83 for the three individual activities and the total score, the percent agreement was ≥ 0.87, and Cronbach’s α was 0.89 with no observed systematic between-rater difference.Conclusions: This study provides evidence for almost perfect inter-rater reliability, excellent internal consistency, and high percent agreement of the Spanish version of the Cumulated Ambulation Score. Due to the strong psychometric properties, and its ease of use, we suggest it be used in Spanish speaking countries to assess early basic mobility status of patients with hip fracture until independence is reached.Implications for rehabilitationThe Spanish version of the Cumulated Ambulation Score is a reliable outcome measure to assess basic mobility of patients with hip fracture.We suggest the Spanish version of the Cumulated Ambulation Score be used in Spanish speaking settings to indicate small changes in basic mobility of patients with hip fracture until an independent level is reached.The Spanish version of the Cumulated Ambulation Score can be used with a high reliability by experienced and inexperienced occupational therapists, corresponding to the already established reliability when used by physicians and physiotherapists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".