Analysis of Evidence‐Based Medicine for Shoulder Instability
Bibliographic record
Abstract
Clinical research has become a major influencing factor in the determination of treatment choice in our society. Outcome data have been requested by third-party payers, patients, and administrators alike. Currently, there are over 10 different scoring systems that have been used to evaluate the efficacy of treatment for shoulder instability. Some of these scoring systems are based on the specific condition of shoulder instability; however, other systems are broadly based to incorporate a spectrum of shoulder conditions. This review summarizes the process of proper development and testing of the scoring systems, discusses their role in clinical research with respect to shoulder instability, and explains the dichotomy of postoperative recurrence of instability and high shoulder scores. The Shoulder Rating Questionnaire (SRQ), Melbourne Instability Shoulder Score (MISS), Western Ontario Shoulder Instability Index (WOSI), Oxford Instability Score (OIS), and Simple Shoulder Test were shown to be reliable for patients with instability. The SRQ, MISS, WOSI, OIS, and American Shoulder and Elbow Surgeons score have all been shown to be largely responsive. There are 2 shoulder scoring systems, the WOSI and the MISS, that we recommend be used to evaluate shoulder instability. The SRQ and OIS were found to be less responsive for patients with instability compared with patients with other shoulder dysfunctions. Other scoring systems lack inter-rater reliability, validity, and/or responsiveness for patients in the instability population. The optimal scoring system for patients with upper extremity problems other than those with shoulder instability has yet to be determined; however, the American Shoulder and Elbow Surgeons score may be considered, because this instrument has been proven to be valid, reliable, and responsive.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".