A Cognitive Autopsy Approach Towards Explaining Diagnostic Failure
Bibliographic record
Abstract
Diagnostic failure has emerged as one of the most significant threats to patient safety. It is important to understand the antecedents of such failures both for clinicians in practice as well is those in training. A consensus has developed in the literature that the majority of failures are due to individual or system factors or some combination of the two. A major source of variance in individual clinical performance is cognitive and affective biases; however, their role in clinical decision making has been difficult to assess partly because they are difficult to investigate experimentally. A significant drawback has been that experimental manipulations appear to confound the assessment of the context surrounding the diagnostic process itself. We conducted an exercise on selected actual cases of diagnostic errors to explore the effect of biases in the 'real world' emergency medicine (EM) context. Thirty anonymized EM cases were analysed in depth through a process of root cause analysis that included an assessment of error-producing conditions (EPCs), knowledge-based errors, and how clinicians were thinking and deciding during each case. A prominent feature of the exercise was the identification of the occurrence of and interaction between specific cognitive and affective biases, through a process called cognitive autopsy. The cases covered a broad range of diagnoses across a wide variety of disciplines. A total of 24 discrete cognitive and affective biases that contributed to misdiagnosis were identified and their incidence recorded. Five to six biases were detected per case, and observed on 168 occasions across the 30 cases. Thirteen EPCs were identified. Knowledge-based errors were rare, occurring in only five definite instances. The ordinal position in which biases appeared in the diagnostic process was recorded. This experiment provides a baseline for investigating and understanding the critical role that biases play in clinical decision making as well as providing a credible explanation for why diagnoses fail.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.161 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".