The challenge of creating new OSCE measures to capture the characteristics of expertise
Bibliographic record
Abstract
PURPOSE: Although expert clinicians approach interviewing in a different manner than novices, OSCE measures have not traditionally been designed to take into account levels of expertise. Creating better OSCE measures requires an understanding of how the interviewing style of experts differs objectively from novices. METHODS: Fourteen clinical clerks, 14 family practice residents and 14 family physicians were videotaped during 2 15-minute standardized patient interviews. Videotapes were reviewed and every utterance coded by type including questions, empathic comments, giving information, summary statements and articulated transitions. Utterances were plotted over time and examined for characteristic patterns related to level of expertise. RESULTS: The mean number of utterances exceeded one every 10 s for all groups. The largest proportion was questions, ranging from 76% of utterances for clerks to 67% for experts. One third of total utterances consisted of a group of 'low frequency' types, including empathic comments, information giving and summary statements. The topic was changed often by all groups. While utterance type over time appeared to show characteristic patterns reflective of expertise, the differences were not robust. Only the pattern of use of summary statements was statistically different between groups (P < 0.05). CONCLUSIONS: Measures that are sensitive to the nature of expertise, including the sequence and organisation of questions, should be used to supplement OSCE checklists that simply count questions. Specifically, information giving, empathic comments and summary statements that occupy a third of expert interviews should be credited. However, while there appear to be patterns of utterances that characterise levels of expertise, in this study these patterns were subtle and not amenable to counting and classification.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".