The Operating Room Black Box: Understanding Adherence to Surgical Checklists
Bibliographic record
Abstract
OBJECTIVE: We report for the first time the use of the Operating Room Black Box (ORBB) to track checklist compliance, engagement, and quality. BACKGROUND: Implementation of operative checklists is associated with improved outcomes. Compliance is difficult to monitor. Most studies report either no assessment of checklist compliance or deployed in-person short-term assessment. The ORBB a novel artificially intelligence-driven data analytic platform affords the opportunity to assess checklist compliance without disrupting surgical workflow. METHODS: This was a retrospective review of prospectively collected ORBB data. Operative cases included elective surgery at a quaternary referral center. Cases were analyzed as prepolicy change (first 9 months) or as a postpolicy change (last 9 months). Measures of checklist compliance, engagement, and quality were assessed. RESULTS: There were 3879 cases that were performed and monitored for checklist compliance between August 15, 2020, and February 20, 2022. The overall scores for compliance, engagement, and quality were 81%, 84%, and 67% respectively. When broken down by phase, the scores for time-out were compliance 100%, engagement 98%, and quality 61%. Scores for the debrief phase were 81% for compliance, 98% for engagement, and 66% for quality. After a hospital policy change, the debrief scores improved significantly (85%; P <0.001 for compliance, 88%; P <0.001 for engagement and 71%; P <0.001 for quality). CONCLUSIONS: ORBB provides the unprecedented ability to assess not only compliance with surgical safety checklists but also engagement and quality. Utilization of this technology allows the assessment of compliance in near real time and to accurately address safety threats that may arise from noncompliance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.084 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".