MétaCan
Menu
Back to cohort
Record W4380762580 · doi:10.4300/jgme-d-22-00783.1

Designing a User-Friendly Context-Specific Assessment Tool for Community-Based Teachers

2023· article· en· W4380762580 on OpenAlexaff
Sanja Kostov, Samantha Horvey, Taryn Wicijowski, Shelley Ross

Bibliographic record

VenueJournal of Graduate Medical Education · 2023
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsRubricSummative assessmentCompetence (human resources)Medical educationComputer scienceContext (archaeology)Formative assessmentPsychologyMedicinePedagogy

Abstract

fetched live from OpenAlex

Graduate medical education (GME) programs must ensure that they are able to collect accurate information about resident competence through assessment tools that are fit for purpose. An assessment form or process is called “fit for purpose” when there is good alignment between the tool (how something is assessed) and the intent (the specific knowledge, skill, or attitude/belief that is being assessed).In our family medicine GME program, we identified that the generic workplace-based summative assessment tool provided to community-based family medicine obstetrics (FMOB) clinical teachers was not fit for purpose. As a result, these teachers were uncertain about program expectations regarding competence, the assessment form was challenging and frustrating to complete, and our program struggled to extract useful assessment data from completed assessment forms. To address this issue, we took a systematic approach to develop a workplace-based assessment tool that was specific to the clinical context of FMOB and user-friendly for clinical teachers.Our goals were to develop: (1) a fit-for-purpose workplace-based assessment tool for community-based FMOB teachers that could be used for accurate assessment without the need for faculty development, and (2) an evidence-guided process for designing similar context-specific tools in the future. We met our first goal through designing the FMOB competence rubric (FMOB-CR) tool. The FMOB-CR is comprised of 2 elements. The first is a rubric containing specific statements and examples of resident performance. The rubric clearly outlines the expected level of competence on a 3-point rating scale (“cause for concern,” “acceptable competence,” “exemplary competence”), organized by the 6 Skill Dimensions of Family Medicine (similar to CanMEDS roles).1 The second element is a simple online assessment form in which FMOB teachers specify their resident's level of competence in each Skill Dimension, guided by the statements and examples in the rubric.The FMOB-CR uses plain language and context-specific clinical examples along with a simple online assessment form to clearly communicate expectations of competence to teachers (and residents) without additional faculty development. This user-friendly tool should make it easier for FMOB clinical teachers to better identify residents who are underperforming, which will allow the program to be more effective in intervening and supporting residents in difficulty.To meet our second goal, we developed an evidence-guided process for designing workplace-based assessment tools that are fit for purpose for specific clinical contexts. Our design process included: assembling a tool development team with relevant expertise, including a resident; consultations with FMOB teachers and residents; environmental scan and review of local assessment forms and assessments used by other FMOB GME programs; a modified Delphi process with local GME program faculty to develop and refine the table of competence statements; a consensus-building process within the research team to revise the statements for the rating scale; and implementation of the assessment rubric with concurrent collection of validity evidence. For both the FMOB-CR and the development process, we collected validity evidence according to Messick's unified concept of validity.2The Table details the validity evidence collected to date. Fifteen FMOB teachers surveyed showed strong agreement across 5 items about the utility of the rubric in allowing them to accurately assess residents (overall M=3.25/4, Likert scale 1=strongly disagree to 4=strongly agree), and across 3 items about the usefulness of the rubrics in helping them to understand program expectations of resident competence (overall M=3.36/4); 14 of 15 teachers preferred the new form to the old one.We hope that the worked example of the FMOB-CR and its development process may serve as a blueprint for other institutions to develop context-specific assessment tools.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.009
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.327
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0070.009
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.084
GPT teacher head0.415
Teacher spread0.331 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Graduate Medical EducationSame topicInnovations in Medical EducationFrench-language works237,207