Measuring the teamwork performance of operating room teams: a systematic review of assessment tools and their measurement properties
Bibliographic record
Abstract
Teamwork is fundamental to surgical patient safety but is inconsistently measured. While many tools have been developed for elective intraoperative situations, it is unclear which is the most robust. This systematic review aimed to identify tools to measure the teamwork of operating room teams. Studies were included if they examined the measurement properties of these tools. PsycINFO, Embase (via OVID), CINAHL, ERIC, Medline and Medline in Process (via OVID) were searched through to May 3, 2019, as were reference lists of included studies and previously published relevant reviews. Retrieved articles were screened and data extracted in duplicate by two independent reviewers. Quality was assessed using the COSMIN checklist. Of the 2121 references identified, 14 studies of six assessment tools were included. Tools were validated across various specialties, mostly in clinical rather than simulated settings. The Observational Teamwork Assessment for Surgery (OTAS) and Operating Theater Team Non-Technical Skills Assessment Tool (NOTECHS) were the most frequently investigated tools. Though acceptable for assessing teamwork, both NOTECHS and OTAS rely on the questionable assumption that the teamwork of a team is equivalent to the sum of individual performances. Future studies may investigate other assessment tools that assess the whole team as the unit of analysis along with the potential of these tools to provide healthcare providers with meaningful feedback in clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".