MétaCan
Menu
Back to cohort
Record W4205794204 · doi:10.1097/acm.0000000000000416

Reflecting Upon Reflection in Diagnostic Reasoning

2014· letter· en· W4205794204 on OpenAlexaffabout
Geoffrey R. Norman, Sandra Monteiro, Jonathan Sherbino

Bibliographic record

VenueAcademic Medicine · 2014
Typeletter
Languageen
FieldMedicine
TopicClinical Reasoning and Diagnostic Skills
Canadian institutionsMcMaster University
Fundersnot available
KeywordsReflection (computer programming)Process (computing)Contrast (vision)CognitionComputer sciencePsychologyTest (biology)Cognitive psychologyArtificial intelligencePsychiatry

Abstract

fetched live from OpenAlex

To the Editor: The recent commentary by Croskerry et al1 discusses the nature of clinical reasoning. The authors critique our experimental study2 contrasting the clinical reasoning of two cohorts of residents who diagnosed a series of written cases, but we disagree with many of their points. In our study, one cohort was told to proceed as quickly as possible without sacrificing accuracy; the other was told to carefully consider all the data. The study was designed to test the hypothesis of dual process theory that diagnostic errors originate from cognitive biases inherent in System 1 (rapid, intuitive) and are corrected by System 2 (slow, analytical) processes. If this is correct, encouraging analytical processing by permitting more time for reflection and systematic inquiry should result in fewer diagnostic errors. We found no support for this hypothesis. Although residents in the slow cohort took an average of 20 seconds longer to complete the case, there was no difference in accuracy. Croskerry et al erroneously state that we assumed that “if decisions are made quickly, they are likely made in the intuitive mode, and therefore making decisions intuitively is a good thing.” In fact, in our study we acknowledge that there are no pure System 1 or 2 tasks. Our study was not intended to contrast Systems 1 and 2; it was designed to examine the effect of varying time and resources available for analytical processing. Croskerry et al also claim that “whether the authors intended it or not, the conclusion most readers will draw from this result … is that diagnostic decisions made faster are more accurate.” We certainly did not intend this conclusion, since the data showed no difference. By contrast, our interpretation was that instructions to be systematic and thorough (and take longer) had no impact on accuracy. We agree with Croskerry et al—it would be silly to “encourage residents … to make speedy diagnoses”; however, we made no such claim. Croskerry et al interpret the results of our first study,4 which did show that accuracy is associated with shorter time, as “residents who knew more and had greater comfort with the material presented were likely to be sufficiently confident to respond more quickly” (and, we might add, more accurately). We completely agree. As we showed,3 (1) self-reported experience with a particular diagnosis related significantly to accuracy, and (2) the disattenuated correlation between diagnostic accuracy and written licensing exam scores was 0.65. This indicates a strong relationship between case-specific knowledge, general clinical knowledge, and accuracy. Errors in diagnosis are more likely to be rectified by conscientious acquisition of relevant knowledge (i.e., clinical experience) than by any attempt to extinguish general cognitive biases and thinking failures. As Graber5 has noted, while the evidence as yet is not strong, the few studies directed at reducing error by explicating cognitive biases have been uniformly negative.5 Conversely, studies directed at improving knowledge application have shown somewhat inconsistent but nevertheless generally positive results.6,7 Finally, Croskerry et al call for more research, both naturalistic and experimental. Certainly this is likely to advance the field more than catchy labels like “creating paralysis by analysis” and “using flawed mental modeling (linear reasoning in complex adaptive systems).”1 Geoffrey Norman, PhD Professor, Department of Clinical Epidemiology and Biostatistics, McMaster University, Hamilton, Ontario, Canada; [email protected] Sandra Monteiro, MSc PhD candidate, Department of Psychology, Neuroscience and Behaviour, McMaster University, Hamilton, Ontario, Canada. Jonathan Sherbino, MD Associate professor, Department of Medicine, McMaster University, Hamilton, Ontario, Canada.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.026
metaresearch head score (Gemma)0.197
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.047
Threshold uncertainty score0.138

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0260.197
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0020.002
Science and technology studies0.0070.014
Scholarly communication0.0110.013
Open science0.0100.006
Research integrity0.0470.055
Insufficient payload (model declined to judge)0.0100.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.084
GPT teacher head0.420
Teacher spread0.337 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2014
Admission routes2
Has abstractyes

Explore more

Same venueAcademic MedicineSame topicClinical Reasoning and Diagnostic SkillsFrench-language works237,207