Variability in Outcome Reporting for Operatively Managed Anterior Glenohumeral Instability: A Systematic Review
Bibliographic record
Abstract
PURPOSE: The purpose of this study was to quantify the degree of variability in outcomes assessed after surgery for anterior shoulder instability in recent high-impact literature. METHODS: Using Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines, an extensive review of the literature during a 5-year period from January 2011 through December 2015 was performed across 6 orthopaedic journals with high impact factors to identify all studies investigating outcomes after anterior shoulder instability. Studies reporting clinical outcomes for patients with anterior glenohumeral instability after surgical treatment with at least 1-year follow-up were included. Several metrics were collected from each manuscript: (1) range of motion (ROM), (2) quantitative strength, (3) physical examination testing, (4) imaging, (5) patient-reported outcomes (PROs), (6) complications (including recurrent instability), (7) patient satisfaction, and (8) return to preinjury level of activity or sport. Variability in outcome measures was then qualitatively assessed. RESULTS: Sixty-eight studies were included for final analysis ranging from Level I to IV evidence. Fifty-nine percent reported ROM, and 18% measured strength. Other clinical exam maneuvers were assessed in 44%, with 40% assessing apprehension. Imaging was used in 62%, including X-rays, magnetic resonance imaging, and computed tomography scans. On average, 2.25 PROs were assessed. In total, 28 different PROs were used to assess outcomes. The 3 most commonly reported PROs were the Rowe scale at 46%, the Western Ontario Shoulder Instability Index at 31%, and the Constant Shoulder Score at 26%. Twenty-five percent included patient satisfaction in their assessment of outcomes. Recurrence was assessed by 59%, and return to preinjury level of activity was reported by 37% of the studies. CONCLUSIONS: There is substantial variability in outcome reporting for high-impact anterior shoulder instability literature with 28 different outcome tools used, making it difficult to compare outcomes between studies. Agreeing upon a uniform measure to assess outcomes would allow for clearer interpretation of the literature as well as the potential to draw conclusions from pooled data. LEVEL OF EVIDENCE: Level IV, systematic review of Level I to IV studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.064 | 0.232 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.014 |
| Bibliometrics | 0.019 | 0.017 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".