MétaCan
Menu
Back to cohort
Record W4320894712 · doi:10.4300/jgme-d-22-00397.1

Program Evaluation Use in Graduate Medical Education

2023· article· en· W4320894712 on OpenAlexafffund
Katherine Moreau, Kaylee Eady

Bibliographic record

VenueJournal of Graduate Medical Education · 2023
Typearticle
Languageen
FieldDecision Sciences
TopicEvaluation and Performance Assessment
Canadian institutionsUniversity of Ottawa
FundersUniversity of Ottawa
KeywordsTimelineProcess (computing)PublicationCurriculumMedical educationProgram evaluationGraduate medical educationComputer sciencePsychologyAccreditationMedicinePolitical sciencePedagogy

Abstract

fetched live from OpenAlex

It is common to complete evaluations of graduate medical education (GME) programs, present them at conferences, publish them in peer-reviewed journals, add them to curricula vitae (CVs), and then move on without using them to enact changes in the programs themselves. Such actions may reflect the reality that many individuals perceive and conduct program evaluations as if they were research.1 While research and program evaluation use similar methods, they have distinct purposes, timelines, audiences, and most notably, intended uses.2 Evaluations of GME programs need to be used to, for example, inform program decisions and modifications, grow program stakeholders' knowledge, stimulate organizational culture changes, or improve the quality of training.1,3,4 They need to be more than intellectual exercises resulting in accomplishments listed on CVs.5 As such, we emphasize that evaluation use is an essential consequence of program evaluation. Those involved in program evaluation should discuss it and maintain its prominence at the onset of every evaluation. We also promote the adage “use-it-or-lose-it” to stress timely program evaluation use. Yet the literature on program evaluation in GME often neglects to discuss use, including how selected evaluation approaches can influence evaluation use.1,6,7 In this article, we explain evaluation use by describing both the use of evaluation findings and process use (ie, changes resulting from engagement in the evaluation process itself).1,8 We also suggest strategies, including evaluation approaches, that faculty can use to increase evaluation use in GME.The 3 categories of use of evaluation findings are instrumental, conceptual, and symbolic. Instrumental use refers to instances where stakeholders use evaluation findings to take direct actions (eg, improvements, changes, terminations) in a program.9 For example, evaluation findings show that residents in a GME program are struggling to complete their research projects. Using the findings, the GME team implements new research training activities to assist residents in the completion of their projects. Conceptual use describes occurrences where stakeholders use evaluation findings to evolve their understandings of a program but do not take direct actions based on these findings.4 For instance, the GME team acknowledges the findings that residents are struggling to complete their research projects. These findings inform their understanding of why residents are not attending academic conferences to present their research. Lastly, symbolic use occurs when stakeholders use the sheer existence of a completed evaluation to comply with reporting requirements or justify a previously made program action.4 For example, the funding university requires the GME program to complete an evaluation to retain funding for residents' research projects. The GME team completes an evaluation and presents the report to the university. Alternatively, before the evaluation, the GME program hired a research assistant to help residents with their research projects and the subsequent evaluation findings are used to justify the hiring of the research assistant. In GME, we emphasize instrumental use, as this form of use leads to actions that can improve programs. However, the use of evaluation findings is typically a short-term consequence of evaluation because these findings are relevant only within a specific and limited timeframe (ie, use-it-or-lose it).On the other hand, process use can have ongoing influence on individuals, programs, and organizations. It recognizes that evaluation processes themselves can affect attitudes, thought processes, and behaviors.10 Process use recognizes stakeholders' learning advancements from their involvement in an evaluation as well as the effects of evaluation processes on program functioning and organizational culture.11 Process use does not require changes to a program or direct actions because of evaluation findings. There are 6 types of process use which we illustrate with examples:When stakeholders are involved in evaluation processes, they enter an evaluation culture and learn how to think and look at things through an evaluative lens. They can also use the knowledge and skills (eg, evaluation knowledge, methodological and facilitation skills) they develop to strengthen their organization's abilities to design, implement, interpret, and use evaluations and thereby build their organization's evaluation capacity. In this sense, process use is valuable throughout and following an evaluation and in various GME settings regardless of the evaluation findings or recommendations.12The Table presents strategies that faculty involved in program evaluation can employ to increase evaluation use.In closing, it is imperative to remember that evaluation use, especially process use, can occur throughout a program evaluation rather than simply at its conclusion.10 Evaluation use can start at the planning stage and continue well beyond a presentation or publication of an evaluation. Program evaluators need a use-it-or-lose-it perspective throughout the evaluation process to maximize improvements to training. This perspective will maintain stakeholders' faith in the value of evaluation, as they witness that evaluation efforts lead to timely, actionable findings and processes. Ultimately, we must embrace evaluation use to ensure that all stakeholders and programs, not only conference attendees, readers of peer-reviewed journals, or our CVs, witness the consequences (both positive and negative) of program evaluation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.039
metaresearch head score (Gemma)0.073
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.853
Threshold uncertainty score0.998

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0390.073
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0020.003
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0010.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.534
GPT teacher head0.616
Teacher spread0.082 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations10
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueJournal of Graduate Medical EducationSame topicEvaluation and Performance AssessmentFrench-language works237,207