MétaCan
Menu
Back to cohort
Record W3096970678 · doi:10.7759/cureus.11363

Rater Training in Medical Education: A Scoping Review

2020· review· en· W3096970678 on OpenAlexaff
Ashley Vergis, Caleb Leung, Reagan Roberston

Bibliographic record

VenueCureus · 2020
Typereview
Languageen
FieldMedicine
TopicClinical Reasoning and Diagnostic Skills
Canadian institutionsUniversity of ManitobaSt. Boniface Hospital
Fundersnot available
KeywordsMedicineMedical educationTraining (meteorology)

Abstract

fetched live from OpenAlex

There is an increasing focus in medical education on trainee evaluation. Often, reliability and other psychometric properties of evaluations fall below expected standards. Rater training, a process whereby raters undergo instruction on how to consistently evaluate trainees and produce reliable and accurate scores, has been suggested to improve rater performance within behavioral sciences. A scoping literature review was undertaken to examine the effect of rater training in medical education and address the question: "Does rater training improve performance attending physician evaluations of medical trainees?" Two independent reviewers searched PubMed®, MEDLINE®, EMBASE™, the Cochrane Library, CINAHL®, ERIC™, and PsycInfo® databases and identified all prospective studies examining the effect of rater training on physician evaluations of medical trainees. Consolidated Standards of Reporting Trials (CONSORT) and Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) checklists were used to assess quality. Fourteen prospective studies met the inclusion criteria. All had heterogeneity in design, type of rater training, and measured outcomes. Pooled analysis was not performed. Four studies examined rater training used to assess technical skills; none identified a positive effect. Ten studies assessed its use to evaluate non-technical skills: six demonstrated no effect, while four showed a positive effect. The overall quality of studies was poor to moderate. Rater training in medical education literature is heterogeneous, limited, and describes minimal improvement on the psychometric properties of trainee evaluations when implemented. Further research is required to assess rater training's efficacy in medical education.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.075
metaresearch head score (Gemma)0.248
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.075
Threshold uncertainty score0.396

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0750.248
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0080.007
Bibliometrics0.0240.026
Science and technology studies0.0020.002
Scholarly communication0.0060.007
Open science0.0040.003
Research integrity0.0050.003
Insufficient payload (model declined to judge)0.0050.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.149
GPT teacher head0.506
Teacher spread0.357 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations22
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueCureusSame topicClinical Reasoning and Diagnostic SkillsFrench-language works237,207