A collective case study of supervision and competence judgments on the inpatient internal medicine ward
Bibliographic record
Abstract
INTRODUCTION: Workplace-based assessment in competency-based medical education employs entrustment-supervision scales to suggest trainee competence. However, clinical supervision involves many factors and entrustment decision-making likely reflects more than trainee competence. We do not fully understand how a supervisor's impression of trainee competence is reflected in their provision of clinical support. We must better understand this relationship to know whether documenting level of supervision truly reflects trainee competence. METHODS: We undertook a collective case study of supervisor-trainee dyads consisting of attending internal medicine physicians and senior residents working on clinical teaching unit inpatient wards. We conducted field observations of typical daily activities and semi-structured interviews. Data was analysed within each dyad and compared across dyads to identify supervisory behaviours, what triggered the behaviours, and how they related to judgments of trainee competence. RESULTS: Ten attending physician-senior resident dyads participated in the study. We identified eight distinct supervisory behaviours. The behaviours were enacted in response to trainee and non-trainee factors. Supervisory behaviours corresponded with varying assessments of trainee competence, even within a dyad. A change in the attending's judgment of the resident's competence did not always correspond with a change in subsequent observable supervisory behaviours. DISCUSSION: There was no consistent relationship between a trigger for supervision, the judgment of trainee competence, and subsequent supervisory behaviour. This has direct implications for entrustment assessments tying competence to supervisory behaviours, because supervision is complex. Workplace-based assessments that capture narrative data including the rationale for supervisory behaviours may lead to deeper insights than numeric entrustment ratings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".