Can Surgeons Adequately Capture Adverse Events Using the Spinal Adverse Events Severity System (SAVES) and OrthoSAVES?
Bibliographic record
Abstract
BACKGROUND: Physicians have consistently shown poor adverse-event reporting practices in the literature and yet they have the clinical acumen to properly stratify and appraise these events. The Spine Adverse Events Severity System (SAVES) and Orthopaedic Surgical Adverse Events Severity System (OrthoSAVES) are standardized assessment tools designed to record adverse events in orthopaedic patients. These tools provide a list of prespecified adverse events for users to choose from-an aid that may improve adverse-event reporting by physicians. QUESTIONS/PURPOSES: The primary objective was to compare surgeons' adverse-event reporting with reporting by independent clinical reviewers using SAVES Version 2 (SAVES V2) and OrthoSAVES in elective orthopaedic procedures. METHOD: This was a 10-week prospective study where SAVES V2 and OrthoSAVES were used by six orthopaedic surgeons and two independent, non-MD clinical reviewers to record adverse events after all elective procedures to the point of patient discharge. Neither surgeons nor reviewers received specific training on adverse-event reporting. Surgeons were aware of the ongoing study, and reported adverse events based on their clinical interactions with the patients. Reviewers recorded adverse events by reviewing clinical notes by surgeons and other healthcare professionals (such as nurses and physiotherapists). Adverse events were graded using the severity-grading system included in SAVES V2 and OrthoSAVES. At discharge, adverse events recorded by surgeons and reviewers were recorded in our database. RESULTS: Adverse-event data for 164 patients were collected (48 patients who had spine surgery, 51 who had hip surgery, 34 who had knee surgery, and 31 who had shoulder surgery). Overall, 99 adverse events were captured by the reviewers, compared with 14 captured by the surgeons (p < 0.001). Surgeons adequately captured major adverse events, but failed to record minor events that were captured by the reviewers. A total of 93 of 99 (94%) adverse events reported by reviewers required only simple or minor treatment and had no long-term adverse effect. Three patients experienced adverse events that resulted in use of invasive or complex treatment that had a temporary adverse effect on outcome. CONCLUSION: Using SAVES V2 and OrthoSAVES, independent reviewers reported more minor adverse events compared with surgeons. The value of third-party reviewers requires further investigation in a detailed cost-benefit analysis. LEVEL OF EVIDENCE: Level II, therapeutic study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".