Obstetrics and Gynecology Modified Delphi Survey for Entrustable Professional Activities: Quantification of Importance, Benchmark Levels, and Roles in Simulation-based Training and Assessment
Bibliographic record
Abstract
Objective Competency-based medical education (CBME) is playing a central role in physicians' training. It focuses on competencies, measured by entrustable professional activities (EPAs). The aim of this survey is threefold for each EPA to (1) quantify the importance for Obstetrics and Gynecology (OBGYN) residency training; (2) set benchmarks; (3) identify the importance of simulation-based training (SBT). Methods The EPAs were defined based on a review of five OBGYN curricula. Two rounds of a modified Delphi via online questionnaire were performed from January to March, 2017. Experts were North American OBGYN program directors. Using a Likert scale, they rated the importance of each EPA for residency training, identified benchmark levels of competence, and roles of simulation. Consensus was defined as ≥80% agreement. Results Item analysis yielded 15 EPAs. Survey response rate was 17.47% (40 out of 229) for part 1 and 6.55% for part 2 (15 out of 229). All experts rated the importance of each EPA for residency training as "moderately important" or "absolutely essential". For benchmarking, experts agreed with a stepwise increase in the level of competence, dependent on residency stage. Two EPAs, "Gynecological Technical Skills & Procedures" and "High-Risk Childbirth", reached consensus (rating 4 or 5) for simulation. Conclusion CBME requires EPAs and benchmarks for each residency stage. Simulation will become a valuable tool in this model. However, experts remain neutral about its role, except for technical skills. An OBGYN curriculum based on predefined EPAs, benchmarks, and adequate assessment tools, including simulation, needs to be further explored for CBME to be successful.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".