Reflecting Upon Reflection in Diagnostic Reasoning
Bibliographic record
Abstract
To the Editor: We are writing regarding the commentary by Croskerry et al1 in the February 2014 issue. The commentary synthesizes important issues confronting investigators interested in the many factors that affect diagnostic reasoning. However, we were surprised by the assertion that the paper-based cases used by Norman et al2 were “so detached from clinical practice as to markedly reduce … validity to the point of making any conclusions extremely tenuous,”1 particularly in view of Croskerry and colleagues’ advocacy of cognitive biases as a principal source of diagnostic error. Theories of cognitive bias stem almost entirely from Tversky and Kahneman’s3 use of paper-based studies with undergraduate psychology students, which were devoid of clinical context. Evidence of cognitive bias in medicine is mostly based on retrospective reviews of adverse events4 and small studies using paper cases.5 A naturalistic setting may add the appearance of ecological validity—as Croskerry et al suggest—but countless variables including limited sampling of context-dependent skills, the common lack of absolute certainty regarding the correct diagnosis, and the empirically established fact that multiple reasoning processes (analytic and nonanalytic) are active any time a judgment is being made,6 confound the ability to infer reasoning from observed behavior in such settings.7 Paper-based cases are not without drawbacks, but enable excellent psychometric properties, provide similar learning outcomes to simulated patient-based cases,8 and correlate with performance in practice.7 Our surprise was heightened given that Croskerry et al provided positive commentary on a study by Schmidt et al5 that used paper-based cases. The difference appears to be that the study by Schmidt et al demonstrated an influence of the availability heuristic. We are less convinced, however, by Croskerry and colleagues’ interpretation that the “deliberate analytical intervention” of reflection is a robust mechanism to optimize diagnostic performance. Schmidt et al effectively demonstrated a benefit of reflection on a subset of cases where availability bias was induced through creation of a deliberate nonanalytic intervention. That does not invalidate the argument that the reasoning processes that create biases exist because they generally offer a useful path towards diagnostic success—not the only path, but a useful path. Deliberate analytic interventions might help in some cases, but can also create detriment in others. We recently compared diagnostic accuracy between participants encouraged to use either first impressions or reflection.9 Our nearly 400 clinician participants (students, residents, and faculty) completed a computer-based assessment using cases drawn from the same collection of paper-based cases used by Schmidt and colleagues.5 Prior to giving their answers, participants were instructed to either “trust [their] sense of familiarity” or engage in structured reflection.9 Those under the reflection condition spent nearly three times longer solving clinical cases (evidence that they followed directions), yet accuracy was identical between the two groups. Consistent with Norman et al,2 we found no evidence that reflection increased accuracy on either straightforward or complex cases. This is not to say that we should give up on reflection, as this exercise may improve learning among novice clinicians. It is just one path to success, however, and it remains unclear how experienced clinicians could be expected to consciously identify “certain situations where we can reasonably and comfortably trust our intuitions, and others where it would be ill advised to use anything other than analytical reasoning.”1 Jonathan S. Ilgen, MD, MCR Assistant professor, Division of Emergency Medicine, University of Washington, School of Medicine, Seattle, Washington; [email protected] Judith L. Bowen, MD Professor, Department of Medicine, Oregon Health & Science University, School of Medicine, Portland, Oregon. Kevin W. Eva, PhD Professor and director of education research and scholarship, Department of Medicine, and senior scientist, Centre for Health Education Scholarship, University of British Columbia, Vancouver, British Columbia, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.053 | 0.356 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.008 | 0.026 |
| Scholarly communication | 0.015 | 0.018 |
| Open science | 0.011 | 0.010 |
| Research integrity | 0.070 | 0.096 |
| Insufficient payload (model declined to judge) | 0.010 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".