MétaCan
Menu
Back to cohort
Record W2091793466 · doi:10.4300/jgme-d-13-00092.1

How Do You Define High-Quality Education Research?

2013· article· en· W2091793466 on OpenAlexaboutno aff
Lalena M. Yarris, Deborah Simpson, Gail M. Sullivan

Bibliographic record

VenueJournal of Graduate Medical Education · 2013
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsnot available
Fundersnot available
KeywordsQuality (philosophy)Data scienceMedical educationComputer scienceMEDLINEMedicinePolitical science

Abstract

fetched live from OpenAlex

Merriam-Webster defines quality as a degree of excellence.1 What is left unstated is how degree and excellence are defined. Does this suggest that the quality of medical education research, like beauty, lies in “the eye of the beholder?” Can we measure quality objectively and consistently or is it subjective and contextual, varying with the type of research question, reviewers' judgments, or quality indices applied? Do these factors capture the aspects of quality that you, our readers, value? We pose these questions for your consideration as you read the following review papers published in this issue of the Journal of Graduate Medical Education (JGME). Locke and colleagues2 reviewed graduate medical education (GME) research papers published in 2011 and selected the 12 articles they considered to be of the greatest importance to internal medicine teachers. With a similar target audience, Eaton et al3 used the Medical Education Research Study Quality Index (MERSQI)4,5 to score internal medicine residency quantitative research papers over a 2-year period. The authors then reviewed the papers ranking in the top 25th percentile for common themes. Examining papers in the surgical education literature published over a decade, Wohlauer and colleagues6 identified common themes and research methods through reviewing the most frequently cited articles in Web of Science, as a surrogate for relevance and quality. Each review aims to identify notable medical education papers for a specific audience and time period, but each takes a different approach. Despite overlapping themes (common topics were simulation, duty hours, resident well-being or distress, resident assessment, and career choices), these 3 reviews achieved different results. Of note, the reviews by Locke et al3 and Eaton et al2 had comparable target audiences, search techniques, and journals reviewed, yet they identified only 2 common papers. The differences may be explained by the use of dissimilar quality criteria, exclusion of qualitative papers for 1 review and only a 50% overlap in review periods. However, the finding that 2 selection processes with a similar aim resulted in almost mutually exclusive results remains striking. The lack of a common definition of quality for medical education research does not stem from a lack of prior efforts to both define and improve the quality of our studies. In addition to the MERSQI, other instruments exist to measure quality in quantitative studies, such as the Best Evidence in Medical Education Global Scale and the Modified Newcastle-Ottawa Scale.7,8 These instruments vary in their (1) incorporation of items that address methodological rigor, (2) reliance on outcome quality based on Kirkpatrick's hierarchy of outcomes of educational interventions, and (3) their association with quality based on a systematic review of method and reporting quality in education research.9–,11 Although methodological rigor is the foundation of quality, attempts to boost quality by focusing on rigor at the expense of other aspects of quality can sometimes diminish the value of the results for consumers. Even the emphasis on outcomes research, a well-intentioned effort to encourage studies that address the highest tier outcomes (patient care or physician behavior outcomes) may result in the unintended consequences of dilution, diminished feasibility, failure to establish a causal link, biased outcome selection, and “teaching to the test.” 12 In addition, we understand that consumers of education research may place value on factors that are not captured by available instruments and that may be neglected by a myopic focus on only the pinnacle of Kirkpatrick's pyramid. The definition of quality for a given product is usually informed by the consumers of that product. Readers of JGME may value elements of quality that are not currently captured by available instruments or methods; we are seeking your input to guide us in future efforts to identify notable medical education papers and help redefine quality in our research.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.036
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.535
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0060.036
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.002
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0000.000
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.091
GPT teacher head0.451
Teacher spread0.359 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2013
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Graduate Medical EducationSame topicInnovations in Medical EducationFrench-language works237,207