International consensus recommendations for outcome measurement in post-stroke arm rehabilitation trials
Bibliographic record
Abstract
BACKGROUND: Existing randomized controlled trials (RCTs) of arm rehabilitation interventions after stroke use a wide range of outcome measures, limiting ability to pool data to determine efficacy. Published recommendations also lack stroke survivor, carer and clinician involvement specifically about perceived relevance and importance of outcomes and measures. AIM: To generate international consensus recommendations for selection of outcome measures for use in future stroke RCTs in arm rehabilitation, considering outcomes important to stroke survivors, carers and clinicians. The recommendations are the Standardizing Measurement in Arm Rehabilitation Trials (SMART) Toolbox. DESIGN: Two-round international e-Delphi Survey and consensus meeting. SETTING: Online and University. POPULATION: Fifty-five researchers and clinicians with expertise in stroke upper limb rehabilitation from 18 countries (e-Delphi); N.=13 researchers and clinicians, N.=2 stroke survivors, N.=1 carer (consensus meeting). METHODS: Using systematically identified outcome measures from published RCTs, we conducted a two-round international e-Delphi Survey with researchers and clinicians to identify the most important measures for inclusion in the toolbox. Measures that achieved ≥60% consensus were categorized using the International Classification of Functioning, Disability and Health Framework (ICF); psychometric properties were ascertained from literature and research resources. At a final consensus meeting, expert stakeholders selected measures for inclusion in the toolbox. RESULTS: e-Delphi participants recommended 28/170 measures for discussion at the final consensus meeting. Expert stakeholders (N.=16) selected the Visual Analogue Scale for pain/0-10 Numeric Pain Rating Scale, dynamometry, Action Research Arm Test, Wolf Motor Function Test, Barthel Index, Motricity Index and Fugl-Meyer Assessment (upper limb section of each), Box and Block Test, Motor Activity Log 14, Nine Hole Peg Test, Functional Independence Measure, EQ-5D, Canadian Occupational Performance Measure and Modified Rankin Scale for inclusion in the toolbox. CONCLUSIONS: The SMART Toolbox provides a refined selection of measures that capture outcomes considered important by stakeholders for each ICF domain. CLINICAL REHABILITATION IMPACT: The toolbox will facilitate data aggregation for efficacy analyses thereby strengthening evidence to inform clinical practice. Clinicians can also use the toolbox to guide selection of measures ensuring a patient-centered focus.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.048 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".