The Operating Room Black Box: Understanding Adherence to Surgical Checklists
Bibliographic record
Abstract
OBJECTIVE: We report for the first time the use of the Operating Room Black Box (ORBB) to track checklist compliance, engagement, and quality. BACKGROUND: Implementation of operative checklists is associated with improved outcomes. Compliance is difficult to monitor. Most studies report either no assessment of checklist compliance or deployed in-person short-term assessment. The ORBB a novel artificially intelligence-driven data analytic platform affords the opportunity to assess checklist compliance without disrupting surgical workflow. METHODS: This was a retrospective review of prospectively collected ORBB data. Operative cases included elective surgery at a quaternary referral center. Cases were analyzed as prepolicy change (first 9 months) or as a postpolicy change (last 9 months). Measures of checklist compliance, engagement, and quality were assessed. RESULTS: There were 3879 cases that were performed and monitored for checklist compliance between August 15, 2020, and February 20, 2022. The overall scores for compliance, engagement, and quality were 81%, 84%, and 67% respectively. When broken down by phase, the scores for time-out were compliance 100%, engagement 98%, and quality 61%. Scores for the debrief phase were 81% for compliance, 98% for engagement, and 66% for quality. After a hospital policy change, the debrief scores improved significantly (85%; P <0.001 for compliance, 88%; P <0.001 for engagement and 71%; P <0.001 for quality). CONCLUSIONS: ORBB provides the unprecedented ability to assess not only compliance with surgical safety checklists but also engagement and quality. Utilization of this technology allows the assessment of compliance in near real time and to accurately address safety threats that may arise from noncompliance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".