Assessing Appropriateness for Shoulder Arthroplasty Using a Shared Decision-Making Process
Bibliographic record
Abstract
Purpose The primary purpose of this study was to validate an appropriateness decision-aid tool as a part of engaging patients with glenohumeral arthritis in their surgical management. The associations between the final decision to have surgery and patient characteristics were examined. Materials and Methods This was an observational study. The demographics, overall health, patient-specific risk profile, expectations, and health-related quality of life were documented. Visual analog scale and the American Shoulder & Elbow Surgeon (ASES) measured pain and functional disability, respectively. Clinical and imaging examination documented clinical findings and extent of degenerative arthritis and cuff tear arthropathy. Appropriateness for arthroplasty surgery was documented by a 5-item Likert response survey and the final decision was documented as ready, not-ready, and would like to further discuss. Results Eighty patients, 38 women (47.5%), mean age: 72(8) participated in the study. The appropriateness decision aid showed excellent discriminate validity (area under the receiver operating characteristic curve value of 0.93) in differentiating between patients who were “ready” and those who were “not-ready” to have surgery. Gender ( P = 0.037), overall health ( P = .024), strength in external rotation ( P = .002), pain severity ( P = .001), ASES score ( P < .0001), and expectations ( P = .024) were contributing factors to the decision to have surgery. Imaging findings did not play a significant role in the final decision to have surgery. Conclusions A 5-item tool showed excellent validity in differentiating patients who were ready to have surgery versus those who were not. Patient's gender, expectations, strength, and self-reported outcomes were important factors in reaching the final decision.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".