Swapping Horses Midstream: Factors Related to Physiciansʼ Changing Their Minds About a Diagnosis
Bibliographic record
Abstract
PURPOSE: Premature closure has been identified as the single most common cause of diagnostic error. This factorial experiment explored which variables exert an unconfounded influence on physicians' diagnostic flexibility (changing their minds about the most likely diagnosis during a clinical case presentation). METHOD: In 2007-2008, 256 practicing physicians viewed a clinically authentic vignette simulating a patient presenting with possible coronary heart disease (CHD) and provided their initial impression midway through the case. At the end, they answered questions about the case, indicated how they would continue their clinical investigation, and made a final diagnosis. The authors used general linear models to determine which patient factors (age, gender, socioeconomic status, race), physician factors (gender, age/experience), and process variables were related to the likelihood of physicians' changing their minds about the most likely diagnosis. RESULTS: Physicians who had less experience, those who named a non-CHD diagnosis as their initial impression, and those who did not ask for information about the patient's prior cardiac disease history were the most likely to change their minds. Participants' certainty in their initial diagnosis, the additional information desired, the diagnostic hypotheses generated, and the follow-up intended were not related to the likelihood of change in diagnostic hypotheses. CONCLUSIONS: Although efforts encouraging physicians to avoid cognitive biases and to reason in a more analytic manner may yield some benefit, this study suggests that experience is a more important determinant of diagnostic flexibility than is the consideration of additional diagnoses or the amount of additional information collected.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.052 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".