The Responsiveness of Patient- Reported Outcome Tools in Shoulder Surgery Is Dependent on the Underlying Pathological Condition
Bibliographic record
Abstract
BACKGROUND: Given the high number of available patient-reported outcome (PRO) tools for patients undergoing shoulder surgery, comparative information is necessary to determine the most relevant forms to incorporate into clinical practice. PURPOSE: To determine the utilization and responsiveness of common PRO tools in studies involving patients undergoing arthroscopic rotator cuff repair or operative management of glenohumeral instability. STUDY DESIGN: Systematic review. METHODS: A systematic review of rotator cuff and instability studies from multiple databases was performed according to PRISMA guidelines. Means and SDs of each PRO tool utilized, study sample sizes, and follow-up durations were collected. The responsiveness of each PRO tool compared with other PRO tools was determined by calculating the effect size and relative efficiency (RE). RESULTS: After a full-text review of 238 rotator cuff articles and 110 instability articles, 81 studies and 29 studies met the criteria for final inclusion, respectively. In the rotator cuff studies, 25 different PRO tools were utilized. The most commonly utilized PRO tools were the Constant (50 studies), visual analog scale (VAS) for pain (44 studies), American Shoulder and Elbow Surgeons (ASES; 39 studies), University of California, Los Angeles (UCLA; 20 studies), and Disabilities of the Arm, Shoulder and Hand (DASH; 13 studies) scores. The ASES score was found to be more responsive than all scores including the Constant (RE, 1.94), VAS for pain (RE, 1.54), UCLA (RE, 1.46), and DASH (RE, 1.35) scores. In the instability studies, 16 different PRO tools were utilized. The most commonly used PRO tools were the ASES (13 studies), Rowe (10 studies), Western Ontario Shoulder Instability Index (WOSI; 8 studies), VAS for pain (7 studies), UCLA (7 studies), and Constant (6 studies) scores. The Rowe score was much more responsive than both the ASES (RE, 22.84) and the Constant (RE, 33.17) scores; however, the ASES score remained more responsive than the Constant (RE, 1.93), VAS for pain (RE, 1.75), and WOSI (RE, 0.97) scores. CONCLUSION: Despite being frequently used in the research community, the Constant score may be less clinically useful as it was less responsive. Additionally, it is a greater burden on the provider because it requires objective strength and range of motion data to be gathered by the clinician. In contrast, the ASES score was highly responsive after rotator cuff repair and requires only subjective patient input. Furthermore, separate PRO scoring methods appear to be necessary for patients undergoing rotator cuff repair and surgery for instability as the instability-specific Rowe score was much more responsive than the ASES score.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.174 | 0.438 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.016 |
| Bibliometrics | 0.011 | 0.016 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".