Quantifying recall bias in surgical safety: a need for a modern approach to morbidity and mortality reviews
Bibliographic record
Abstract
Background: Despite recent investments into reducing errors and adverse events in health care, methods for quality improvement in surgery are outdated and ineffective. Most current efforts in this field are centred around morbidity and mortality conferences (MMCs), which have remained unchanged for over 100 years. The present study aimed to quantify the recall bias associated with details from surgical cases. Methods: We gathered immediate postoperative questionnaires from 1 surgeon, 1 fellow and 11 trainees following 25 routine surgical cases. Information elicited included their perceived level of concentration, mental preparedness and assessment of whether the procedure deviated from its expected course, including any intraoperative adverse events. We readministered the questionnaire 7−9 days later to assess participants’ ability to recall important aspects of the procedure. Results: After 1 week, members of the surgical team were universally inaccurate in their recollection of even major details from the operating room. Although most participants felt mentally prepared and perceived no issues with concentration during the case, all participants misclassified operations as having been performed with or without adverse events in almost every included case. Conclusion: Our findings show that recall bias regarding surgical safety events is exceedingly common. This likely has a major impact on the integrity of data presented at MMCs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".