MétaCan
Menu
← Back to cohort
Record W3208344337 · doi:10.1097/acm.0000000000004282

A Question of Scale? Comparison of Generalizability in Ottawa and Chen Scales When Used to Formulate Ad Hoc Entrustment Decisions for the Core EPAs

2021· article· en· W3208344337 on OpenAlexaboutno aff
Michael S. Ryan, Rebecca Khamishon, Alicia Richards, Robert A. Perera, Adam Garber, Sally A. Santen

Bibliographic record

VenueAcademic Medicine · 2021
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsnot available
Fundersnot available
KeywordsGeneralizability theoryScale (ratio)Formative assessmentReliability (semiconductor)Medical educationPsychologyMedicineMathematics educationGeography

Abstract

fetched live from OpenAlex

Purpose: Assessment of the Core Entrustable Professional Activities (Core EPAs) is based on observations of supervisors throughout a medical student’s progression toward entrustment. 1,2 In our previous work, we examined performance of the Ottawa Clinic Assessment Tool (Ottawa) 3 when used to measure medical student performance of the Core EPAs in the workplace setting. 4 The findings from that study demonstrated poor generalizability of the Ottawa scale. 4 The primary purpose of the present study was to compare performance of the Ottawa scale with a second scale—the undergraduate medical education (UME) supervisory scale proposed by Chen and colleagues (Chen) 5 in terms of reliability and generalizability. A secondary aim was to determine the impact of frequent assessors on the validity and reliability of the data. Methods: For the 2019–2020 academic year, the Virginia Commonwealth University School of Medicine modified a previously described, 4 student-initiated, workplace-based assessment (WBA) system developed to provide formative feedback for the Core EPAs across clerkships. The WBA scored students’ performance using both the modified Ottawa and the modified Chen scales. Generalizability and decision studies were performed to determine the reliability of each scale. Secondary analysis explored whether faculty who frequently assess the EPAs demonstrated better reliability. Results: A total of 923 raters completed 7,277 WBAs on 208 medical students across all clerkships. Using Ottawa, variability attributable to the student ranged from 0.8% to 6.5%. For Chen, variability attributable to the student ranged from 1.8% to 7.1%. These findings indicate that the majority of the variation for EPA ratings was due to the rater (42.8%–61.3%) and other unexplained factors. A range of 28 to 127 assessments were required to obtain a Phi coefficient of 0.70. For 2 EPAs, using only faculty who frequently assess the EPA improved generalizability—requiring only 5 and 13 assessments for the Chen scale. Discussion: Both the Ottawa and Chen scales performed poorly in terms of variance attributed to the learner. The frequent assessor model seemed to increase variance attributed to the learner for the Chen scale in only 2 Core EPAs. Overall, these findings were similar to our previous study 4 involving only the Ottawa scale. Based on these findings in conjunction with prior evidence, we suggest that the root cause analysis for challenges associated with WBAs for the Core EPAs involves the lack of a true competency-based curriculum in UME; the choice of scale alone does not appear to impact performance. Significance: This study adds to the emerging literature around WBAs specific to the Core EPAs in UME. Based on these findings as well as those from our prior work, we feel there is a need to reconsider the workflow, scale, and investment of learners and faculty to best assess the Core EPAs in the UME setting. Acknowledgments: The authors would like to thank Joel Browning (former director of academic information systems at Virginia Commonwealth University School of Medicine [VCU-SOM]), who worked with Brie Dubinsky, MS, to develop the workplace-based assessment system described in this manuscript. In addition, the authors would like to thank Yoon Soo Park, PhD, for his assistance with statistical methodology.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.123
metaresearch head score (Gemma)0.325
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.877
Threshold uncertainty score0.653

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1230.325
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.004
Bibliometrics0.0050.004
Science and technology studies0.0010.004
Scholarly communication0.0030.004
Open science0.0020.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.095
GPT teacher head0.437
Teacher spread0.342 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueAcademic Medicine→Same topicInnovations in Medical Education→French-language works237,207→