In Reply to Norman et al and to Ilgen et al
Bibliographic record
Abstract
Norman et al’s response evades our point that their experimental conditions did not distinguish System 1 from System 2 processes. Again, in the time available to the res pondents in their study1 and its predecessor,2 it was quite possible to make an analytical response to a written clinical case containing the relevant data necessary to make a correct diagnosis. Therefore, speed of response alone did not distinguish individuals processing intuitively from those processing analytically. Any discussion of System 1 and System 2 in their papers is digressive, since there is no evidence that they are triggering System 1 processes in the cases they used. The variability in accuracy they observed could be explained by the extent of the respondent’s knowledge, and perhaps by their particular level of confidence. The information presented to respondents in these studies presumably had a high level of consistency, as all the data necessary to make a correct diag nosis would have to be present, and, when data consistency is high, diagnostic accuracy is associated with higher levels of certainty.3 They also state that “Errors in diagnosis are more likely to be rectified by con scientious acquisition of relevant knowledge (i.e. clinical experience) than by any attempt to extinguish general cognitive biases and thinking failures.” We agree; it is difficult to conceive of someone being a good diagnostician without knowledge and experience. However, it is misleading to present it as a choice. The question isn’t whether excellent content knowledge is better than de-biasing, or even whether System 2 is better than System 1. The question should be: Given the same content knowledge in the same real-world clinical scenario, does awareness of one’s potential biases and flawed decision-making habits, and use of reflective intelligence or appropriate “mindware”4 to make the necessary adjustments, improve performance? Finally, the statement that “the few studies directed at reducing error by explicating cognitive biases have been uniformly negative” is simply inaccurate. Graber et al’s review of 140 studies aimed at cognitive interventions to improve clinical reasoning and decision-making, which included reflective practice and active metacognitive review, found a number that showed positive, beneficial effects.5 To condemn the potential value of such work is premature. Ilgen et al misrepresent our position on “paper-based cases.” We were commenting on the unrepresentative nature of the conditions under which attempts were made to investigate clinical decision making (CDM).1,2 Written cases certainly have a role in the study of CDM but with the strong caveat that their use is as close as reasonably possible to the real clinical world. The study that we favourably commented upon6 was clearly in this category. Consideration of external and ecological validity is critically important in the interpretation of experimental CDM studies. Their comment that we might be biased ourselves is quite correct. It is widely appreciated that while scientists are ostensibly committed to the scientific principle of objectivity, they may ignore findings that contradict what they already believe—a relatively recent review of suggested educational strategies to promote clinical diagnostic reasoning omitted any discussion of cognitive bias.7 Bias is a normal operating characteristic of the human brain,8 and it is time to forego an ostrich-like attitude towards it. Their final point in which they challenge the value of reflection will be anathema to many who value the role of reflection in achieving quality of thought. We strongly recommend the broad perspective offered in Epstein’s classic paper,9 which soundly explores the value of reflection and mindfulness in clinical practice. Pat Croskerry, MD, PhD Professor and director, Critical Thinking Program, Division of Medical Education, Faculty of Medicine, Dalhousie University, Halifax, Nova Scotia, Canada; e-mail: [email protected] David A. Petrie, MD Professor of emergency medicine and professor, Department of Emergency Medicine, Faculty of Medicine, Dalhousie University, and chief, Capital District Health Authority Department of Emergency Medicine, Halifax, Nova Scotia, Canada. James B. Reilly, MD, MS Associate director, Internal Medicine Residency, Allegheny General Hospital, Western Pennsylvania Hospital Educational Consortium, Pittsburgh, Pennsylvania, and assistant professor of medicine, Temple University School of Medicine, Philadelphia, Pennsylvania. Gordon Tait, PhD Assistant professor, Departments of Surgery and Anesthesia, and staff scientist, Department of Anesthesia, Toronto General Hospital, University Health Network, Faculty of Medicine, University of Toronto, Toronto, Ontario, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.098 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.005 | 0.007 |
| Scholarly communication | 0.007 | 0.010 |
| Open science | 0.007 | 0.004 |
| Research integrity | 0.062 | 0.082 |
| Insufficient payload (model declined to judge) | 0.011 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".