A Systematic Review of the Psychometric Properties of Patient-Reported Outcome Instruments for Use in Patients With Rotator Cuff Disease
Bibliographic record
Abstract
BACKGROUND: Many patient-reported outcome instruments (or questionnaires) have been developed for use in patients with rotator cuff disease. Before an instrument is implemented, its psychometric properties should be carefully assessed, and the methodological quality of papers that investigate a psychometric component of an instrument must be carefully evaluated. Together, the psychometric evidence and the methodological quality can then be used to arrive at an estimate of an instrument's quality. PURPOSE: To identify patient-reported outcome instruments used in patients with rotator cuff disease and to critically appraise and summarize their psychometric properties to guide researchers and clinicians in using high-quality patient-reported outcome instruments in this population. STUDY DESIGN: Systematic review. METHODS: Systematic literature searches were performed to find English-language articles concerning the development or evaluation of a psychometric property of a patient-reported outcome instrument for use in patients with rotator cuff disease. Methodological quality and psychometric evidence were critically appraised and summarized through 2 standardized sets of criteria. RESULTS: A total of 1881 articles evaluating 39 instruments were found per the search strategy, of which 73 articles evaluating 16 instruments were included in this study. The Constant-Murley score, the DASH (Disability of the Arm, Shoulder, and Hand), and the Shoulder Pain and Disability Index were the 3 most frequently evaluated instruments. In contrast, the psychometric properties of the Korean Shoulder Scoring System, Shoulder Activity Level, Subjective Shoulder Value, and Western Ontario Osteoarthritis Shoulder index were evaluated by only 1 study each. The Western Ontario Rotator Cuff Index was found to have the best overall quality of psychometric properties per the established criteria, with positive evidence found in internal consistency, reliability, content validity, hypothesis testing, and responsiveness. The DASH, Shoulder Pain and Disability Index, and Simple Shoulder Test had good evidence in support of internal consistency, reliability, structural validity, hypothesis testing, and responsiveness. Inadequate methodological quality was found across many studies, particularly in internal consistency, reliability, measurement error, hypothesis testing, and responsiveness. CONCLUSION: More high-quality methodological studies should be performed to assess the properties in all identified instruments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.006 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".