Reflecting Upon Reflection in Diagnostic Reasoning
Notice bibliographique
Résumé
To the Editor: We are writing regarding the commentary by Croskerry et al1 in the February 2014 issue. The commentary synthesizes important issues confronting investigators interested in the many factors that affect diagnostic reasoning. However, we were surprised by the assertion that the paper-based cases used by Norman et al2 were “so detached from clinical practice as to markedly reduce … validity to the point of making any conclusions extremely tenuous,”1 particularly in view of Croskerry and colleagues’ advocacy of cognitive biases as a principal source of diagnostic error. Theories of cognitive bias stem almost entirely from Tversky and Kahneman’s3 use of paper-based studies with undergraduate psychology students, which were devoid of clinical context. Evidence of cognitive bias in medicine is mostly based on retrospective reviews of adverse events4 and small studies using paper cases.5 A naturalistic setting may add the appearance of ecological validity—as Croskerry et al suggest—but countless variables including limited sampling of context-dependent skills, the common lack of absolute certainty regarding the correct diagnosis, and the empirically established fact that multiple reasoning processes (analytic and nonanalytic) are active any time a judgment is being made,6 confound the ability to infer reasoning from observed behavior in such settings.7 Paper-based cases are not without drawbacks, but enable excellent psychometric properties, provide similar learning outcomes to simulated patient-based cases,8 and correlate with performance in practice.7 Our surprise was heightened given that Croskerry et al provided positive commentary on a study by Schmidt et al5 that used paper-based cases. The difference appears to be that the study by Schmidt et al demonstrated an influence of the availability heuristic. We are less convinced, however, by Croskerry and colleagues’ interpretation that the “deliberate analytical intervention” of reflection is a robust mechanism to optimize diagnostic performance. Schmidt et al effectively demonstrated a benefit of reflection on a subset of cases where availability bias was induced through creation of a deliberate nonanalytic intervention. That does not invalidate the argument that the reasoning processes that create biases exist because they generally offer a useful path towards diagnostic success—not the only path, but a useful path. Deliberate analytic interventions might help in some cases, but can also create detriment in others. We recently compared diagnostic accuracy between participants encouraged to use either first impressions or reflection.9 Our nearly 400 clinician participants (students, residents, and faculty) completed a computer-based assessment using cases drawn from the same collection of paper-based cases used by Schmidt and colleagues.5 Prior to giving their answers, participants were instructed to either “trust [their] sense of familiarity” or engage in structured reflection.9 Those under the reflection condition spent nearly three times longer solving clinical cases (evidence that they followed directions), yet accuracy was identical between the two groups. Consistent with Norman et al,2 we found no evidence that reflection increased accuracy on either straightforward or complex cases. This is not to say that we should give up on reflection, as this exercise may improve learning among novice clinicians. It is just one path to success, however, and it remains unclear how experienced clinicians could be expected to consciously identify “certain situations where we can reasonably and comfortably trust our intuitions, and others where it would be ill advised to use anything other than analytical reasoning.”1 Jonathan S. Ilgen, MD, MCR Assistant professor, Division of Emergency Medicine, University of Washington, School of Medicine, Seattle, Washington; [email protected] Judith L. Bowen, MD Professor, Department of Medicine, Oregon Health & Science University, School of Medicine, Portland, Oregon. Kevin W. Eva, PhD Professor and director of education research and scholarship, Department of Medicine, and senior scientist, Centre for Health Education Scholarship, University of British Columbia, Vancouver, British Columbia, Canada.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,053 | 0,356 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,003 | 0,003 |
| Bibliométrie | 0,003 | 0,002 |
| Études des sciences et des technologies | 0,008 | 0,026 |
| Communication savante | 0,015 | 0,018 |
| Science ouverte | 0,011 | 0,010 |
| Intégrité de la recherche | 0,070 | 0,096 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,010 | 0,006 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».