In the Real World, Faster Diagnoses Are Not Necessarily More Accurate
Bibliographic record
Abstract
In Reply to Reilly and Von Feldt: We thank Drs. Reilly and Von Feldt for their interest in our study. However, their main concern seems to be based on a misunderstanding of the methodology. They write that “their conclusion appears valid in the context of a multiple-choice examination but … is poorly applicable to real-life clinical situations.” In fact, as we described, the cases we used were long, written ones (about 400 words) requiring the participant physician to type in a free-text diagnostic response. Their commentary does raise two important questions: (1) What is the relationship between performance on written test cases and performance in “the real world”? and (2) Do “cognitive bias effects” exert any greater influence in “the real world” than in the written case? The first question can be easily answered. Studies done by the Medical Council of Canada1 examined the validity of the two-part Medical Council licensing examinations in predicting complaints to a regulatory agency. Part 1, the written, mostly multiple-choice examination, had a substantially greater relationship to complaints than measures of problem-solving derived from the Part 2 OSCE performance examination. Moreover, the written case-based clinical decision-making component of the Part 1 examination, which resembles our methodology, had the best predictive validity. As to the second question, the empirical evidence for virtually all cognitive biases is based entirely on written materials,2 often with multiple-choice questions, administered to undergraduate psychology students. To argue that written materials cannot identify cognitive biases negates the entire evidential basis of the cognitive bias literature. Moreover, the argument that these biases are more frequent in the clinical situation is not substantiated by the studies cited by Reilly and Von Feldt, which do not make any comparison to written cases, and Zwaan and colleagues’3 study counts “mistakes,” “violations,” “slips,” and “lapses,” which do not equate to cognitive biases. These are not inconsequential issues. If one assumes that the majority of diagnostic errors arise from cognitive biases that originate in System I reasoning, then it would be appropriate to devise educational interventions that, in Reilly and Von Feldt’s words, “encourage our learners to … be mindful of [System I’s] potential to introduce bias to diagnostic decision making.” In our view, such an assumption would also mean that learners should be encouraged to slow down and avoid System I reasoning. But the evidence we have presented shows that rapid diagnoses are more, not less, accurate on average, so an intervention to universally discourage speed would likely be ineffective and wasteful. Jonathan Sherbino, MD Associate professor, Emergency Medicine, McMaster University Faculty of Health Sciences, Hamilton, Ontario, Canada. Geoffrey R. Norman, PhD Professor, Clinical Epidemiology and Biostatistics, McMaster University Faculty of Health Sciences, Hamilton, Ontario, Canada; [email protected]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.199 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.010 |
| Scholarly communication | 0.005 | 0.014 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.026 | 0.047 |
| Insufficient payload (model declined to judge) | 0.011 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".