Analysis of Evidence‐Based Medicine for Shoulder Instability
Bibliographic record
Abstract
Clinical research has become a major influencing factor in the determination of treatment choice in our society. Outcome data have been requested by third-party payers, patients, and administrators alike. Currently, there are over 10 different scoring systems that have been used to evaluate the efficacy of treatment for shoulder instability. Some of these scoring systems are based on the specific condition of shoulder instability; however, other systems are broadly based to incorporate a spectrum of shoulder conditions. This review summarizes the process of proper development and testing of the scoring systems, discusses their role in clinical research with respect to shoulder instability, and explains the dichotomy of postoperative recurrence of instability and high shoulder scores. The Shoulder Rating Questionnaire (SRQ), Melbourne Instability Shoulder Score (MISS), Western Ontario Shoulder Instability Index (WOSI), Oxford Instability Score (OIS), and Simple Shoulder Test were shown to be reliable for patients with instability. The SRQ, MISS, WOSI, OIS, and American Shoulder and Elbow Surgeons score have all been shown to be largely responsive. There are 2 shoulder scoring systems, the WOSI and the MISS, that we recommend be used to evaluate shoulder instability. The SRQ and OIS were found to be less responsive for patients with instability compared with patients with other shoulder dysfunctions. Other scoring systems lack inter-rater reliability, validity, and/or responsiveness for patients in the instability population. The optimal scoring system for patients with upper extremity problems other than those with shoulder instability has yet to be determined; however, the American Shoulder and Elbow Surgeons score may be considered, because this instrument has been proven to be valid, reliable, and responsive.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.130 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.007 | 0.007 |
| Bibliometrics | 0.015 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.007 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.011 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".