Family Medicine Residents’ Performance with Detected Versus Undetected Simulated Patients Posing as Problem Drinkers
Bibliographic record
Abstract
BACKGROUND: Simulated patients are commonly used to evaluate medical trainees. Unannounced simulated patients provide an accurate measure of physician performance. PURPOSE: To determine the effects of detection of SPs on physician performance, and identify factors leading to detection. METHODS: Fixty-six family medicine residents were each visited by two unannounced simulated patients presenting with alcohol-induced hypertension or insomnia. Residents were then surveyed on their detection of SPs. RESULTS: SPs were detected on 45 out of 104 visits. Inner city clinics had higher detection rates than middle class clinics. Residents' checklist and global rating scores were substantially higher on detected than undetected visits, for both between-subject and within-subject comparisons. The most common reasons for detection concerned SP demographics and behaviour; the SP "did not act like a drinker" and was of a different social class than the typical clinic patient. CONCLUSIONS: Multi-clinic studies involving residents experienced with SPs should ensure that the SP role and behavior conform to physician expectations and the demographics of the clinic. SP station testing does not accurately reflect physicians' actual clinical behavior and should not be relied on as the primary method of evaluation. The study also suggests that physicians' poor performance in identifying and managing alcohol problems is not entirely due to lack of skill, as they demonstrated greater clinical skills when they became aware that they were being evaluated. Physicians' clinical priorities, sense of responsibility and other attitudinal determinants of their behavior should be addressed when training physicians on the management of alcohol problems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".