MétaCan
Menu
Back to cohort
Record W2136442190 · doi:10.1080/01421590500136113

Creating a reliable and valid blueprint for the internal medicine clerkship evaluation

2005· article· en· W2136442190 on OpenAlexaffabout
Kevin McLaughlin, Jane B Lemaire, Sylvain Coderre

Bibliographic record

VenueMedical Teacher · 2005
Typearticle
Languageen
FieldHealth Professions
TopicPatient Satisfaction in Healthcare
Canadian institutionsUniversity of Calgary
Fundersnot available
KeywordsBlueprintClinical clerkshipMedicineObjective structured clinical examinationCohen's kappaPresentation (obstetrics)Medical educationMedical physicsFamily medicinePsychologyCurriculumStatisticsSurgeryMathematicsPedagogy

Abstract

fetched live from OpenAlex

The objective of this study was to design an examination blueprint for the Internal Medicine clerkship rotation that is congruent with both the learning objectives and delivered learning experiences and reflects the perceived importance of clinical presentations from both the students' and clinicians' perspectives. In this cross-sectional study 11 specialists in General Internal Medicine (GIM) and 11 clinical clerks at the University of Calgary were asked to score each of the 47 clinical presentations in the Internal Medicine clerkship rotation for 'impact' and 'frequency'. These attributes were used to provide an estimate of the relative importance of each clinical presentation. Statistical tests used were the Pearson's correlation coefficient and the Kappa statistic. Multi-attribute utility theory was applied to assess the best way of combining the variables of 'impact' and 'frequency'. The correlation between clerks and GIM specialists was 0.85 for the impact score and 0.86 for the frequency score (p < 0.001 for both). Corresponding Kappa values were 0.71 and 0.82, respectively (p < 0.001 for both). Combining impact and frequency as a multiplicative function produced a distribution that was positively skewed towards common, high impact presentations such as chest pain. We have created an examination blueprint that provides a realistic and objective measure of the relative importance of clinical presentations. Such a blueprint provides both face validity and content validity to the evaluation process.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.008
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.428
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0070.008
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0120.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.238
GPT teacher head0.527
Teacher spread0.289 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations21
Published2005
Admission routes2
Has abstractyes

Explore more

Same venueMedical TeacherSame topicPatient Satisfaction in HealthcareFrench-language works237,207