Transforming Team Performance Through Reimplementation of the Surgical Safety Checklist
Bibliographic record
Abstract
Importance: Patient safety interventions, like the World Health Organization Surgical Safety Checklist, require effective implementation strategies to achieve meaningful results. Institutions with underperforming checklists require evidence-based guidance for reimplementing these practices to maximize their impact on patient safety. Objective: To assess the ability of a comprehensive system of safety checklist reimplementation to change behavior, enhance safety culture, and improve outcomes for surgical patients. Design, Setting, and Participants: This prospective type 2 hybrid implementation-effectiveness study took place at 2 large academic referral centers in Singapore. All operations performed at either hospital were eligible for observation. Surveys were distributed to all operating room staff. Intervention: The study team developed a comprehensive surgical safety checklist reimplementation package based on the Exploration, Preparation, Implementation, Sustainment framework. Best practices from implementation science and human factors engineering were combined to redesign the checklist. The revised instrument was reimplemented in November 2021. Main Outcomes and Measures: Implementation outcomes included penetration and fidelity. The primary effectiveness outcome was team performance, assessed by trained observers using the Oxford Non-Technical Skills (NOTECH) system before and after reimplementation. The Agency for Healthcare Research and Quality Hospital Survey on Patient Safety Culture was used to assess safety culture and observers tracked device-related interruptions (DRIs). Patient safety events, near-miss events, 30-day mortality, and serious complications were tracked for exploratory analyses. Results: Observers captured 252 cases (161 baseline and 91 end point). Penetration of the checklist was excellent at both time points, but there were significant improvements in all measures of fidelity after reimplementation. Mean NOTECHS scores increased from 37.1 to 42.4 points (4.3 point adjusted increase; 95% CI, 2.9-5.7; P < .001). DRIs decreased by 86.5% (95% CI, -22.1% to -97.8%; P = .03). Significant improvements were noted in 9 of 12 composite areas on culture of safety surveys. Exploratory analyses suggested reductions in patient safety events, mortality, and serious complications. Conclusions and Relevance: Comprehensive reimplementation of an established checklist intervention can meaningfully improve team behavior, safety culture, patient safety, and patient outcomes. Future efforts will expand the reach of this system by testing a structured guidebook coupled with light-touch implementation guidance in a variety of settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".