MétaCan
Menu
← Back to cohort
Record W4391886477 · doi:10.1093/jcag/gwad061.084

A84 COMPETENCY BASED MEDICAL EDUCATION: O-SCORE CHARACTERISTICS OF PROCEDURAL AND COGNITIVE ASSESSMENTS IN GASTROENTEROLOGY RESIDENCY TRAINING

2024· article· en· W4391886477 on OpenAlexaffabout
Judith Cooper, M Gozdzik, Jason A. Silverman, Karen I. Kroeker

Bibliographic record

VenueJournal of the Canadian Association of Gastroenterology · 2024
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsResidency trainingMedicineCognitionMedical educationPsychologyInternal medicineContinuing education

Abstract

fetched live from OpenAlex

Abstract Background Competency based medical education has become the new standard for medical education which shifts the focus of training toward a competency, rather than time-based in framework known in Canada as ‘Competence by Design’ (CBD). CBD assesses a physician trainee’s ability to demonstrate competence in CanMEDS roles via entrustable professional activities (EPAs). EPAs utilize the O-SCORE as the metric for assessing competence. This score was developed and validated for surgical/procedural subspecialties; however, CBD currently coopts this scale for both procedural and non-procedural (cognitive) EPAs. Assessor expertise has also been shown to have an important role in performance assessments, but has not been studied in the context of CBD. Aims Our study aims to assess for differences in O-SCORE utilization between cognitive and procedural EPAs, and whether assessor characteristics are associated with trends in assessment. Methods Anonymized data for all Adult GI subspecialty EPAs completed from Jun 2019 to Jan 2023 at the University of Alberta was obtained. Evaluator sex, clinical vs academic practice, advanced training expertise, and EPA score was extracted. Locally a score of 5 denotes competence, while a 1-3 indicates competence was not yet achieved. A score of 4 may be accepted as evidence of competence (neutral score), at the discretion of the local competency committee. Data was analyzed via T-tests and ANOVA with post hoc Games-Howell testing with 95% confidence intervals (CI). A p-value of ampersand:003C0.05 was significant. Results 2264 EPAs were assessed including 1385 cognitive and 879 procedural EPAs. The number of EPAs completed by evaluators ranged from 11 to 165 with a mean of 60 (standard deviation: 40). Results of O-SCORE usage is summarized in Figure 1A-B. The majority of EPAs indicate competence, with 20-25% neutral, and ampersand:003C10% did not achieve competence. Less than one of third of evaluators utilized a score of 1 or 2 across all EPAs, and zero evaluators utilized a score of 1 for cognitive EPAs. Most commonly evaluators to utilized 3/5 options of the O-SCORE. Separated by EPA type, it was most common to utilize 2/5 and 4/5 options for cognitive and procedural EPAs respectively. Results of demographic comparisons are outlined if Figure 1C-E. Male and clinical evaluators submitted higher scores on average. Hepatologists submitted higher scores than all other advanced training areas for total, cognitive, and procedural EPAs. Conclusions Across total, cognitive, and procedural EPAs there are low rates in the utilization of the whole O-SCORE scale, and our study highlights a discrepancy between procedural and cognitive EPAs. In addition, there small but significant differences in the mean EPAs score awarded between different evaluator demographics (male, clinical, hepatologists providing higher scores). Figure 1. A) Number and proportion of Entrustable Professional Activities (EPA) stratified by type and competence evaluation. B) Number and proportion of Entrustable Professional Activities (EPA) stratified by type with scored 1-5 and percent (%) of staff utilizing each score stratified by EPA type C) Number, mean, and mean difference of Entrustable Professional Activities (EPA) stratified by evaluator sex and EPA type. D) Number, mean, and mean difference of Entrustable Professional Activities (EPA) stratified by evaluator academic vs clinical status and EPA type. E) Number, mean, and mean difference of Entrustable Professional Activities (EPA) stratified by evaluator advanced training and EPA type. CI: Confidence interval; SD: standard deviation; *: pampersand:003C0.05 Funding Agencies None

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.022
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.017
Threshold uncertainty score0.034

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.022
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0020.002
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.002
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.014
GPT teacher head0.313
Teacher spread0.299 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes2
Has abstractyes

Explore more

Same venueJournal of the Canadian Association of Gastroenterology→Same topicInnovations in Medical Education→French-language works237,207→