An Entrustable Professional Activities Model for Assessment of Undergraduate Competence in Anesthesia and Surgery: Performance of the Assessment Scheme and the Impact of Assessment Timing and Variation in Structure of Teaching Activities on Student Outcomes
Bibliographic record
Abstract
An entrustable professional activity (EPA) model was used to assess the anesthesia and surgery competence of year 4 students during elective neutering procedures over 3 academic years (cohort A, cohort B, and cohort C). Two competence thresholds were defined by an expert panel, the minimum acceptable standard (MAS) and the standard expected at the start of final-year rotations (SFR). The assessment scheme performed as expected, and the median level of supervision achieved by students either matched or exceeded the SFR for all EPAs except one, which matched the MAS. Semester of assessment was associated with student performance, with more students in semester 2 achieving the SFR. In the EPAs assessing pain management, documentation, and patient discharge, cohort A was associated with reduced student performance; this could be explained by changes in the delivery of teaching that enhanced performance in subsequent cohorts (academic years). For all EPAs combined and for EPAs 3, 5, 6, 8, and 9, student performance at the SFR was associated with academic year. For all EPAs combined and EPAs 3, 8, and 9, there was a reduction in the proportion of students achieving the SFR threshold in each successive year. At the MAS, the only association for all EPAs combined was with cohort C. This progressive reduction in performance may have been related to the negative effect of decreased time spent at the neutering clinic and loss of feedback opportunities outweighing the positive effects of increased staff:student ratio and improvements in the preparative phases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".