Is There a Problem With Evidence in Health Professions Education?
Bibliographic record
Abstract
ABSTRACT: What constitutes evidence, what value evidence has, and how the needs of knowledge producers and those who consume this knowledge might be better aligned are questions that continue to challenge the health sciences. In health professions education (HPE), debates on these questions have ebbed and flowed with little sense of resolution or progress. In this article, the authors explore whether there is a problem with evidence in HPE using thought experiments anchored in Argyris' learning loops framework.From a single-loop perspective ("How are we doing?"), there may be many problems with evidence in HPE, but little is known about how research evidence is being used in practice and policy. A double-loop perspective ("Could we do better?") suggests expectations of knowledge producers and knowledge consumers might be too high, which suggests more system-wide approaches to evidence-informed practice in HPE are needed. A triple-loop perspective ("Are we asking the right questions?") highlights misalignments between the dynamics of research and decision-making, such that scholarly inquiry may be better approached as a way of advancing broader conversations, rather than contributing to specific decision-making processes.The authors ask knowledge producers and consumers to be more attentive to the translation from knowledge to evidence. They also argue for more systematic tracking and audit of how research knowledge is used as evidence. Given that research does not always have to serve practical purposes or address the problems of a particular program or institution, the relationship between knowledge and evidence should be understood in terms of changing conversations and influencing decisions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.643 | 0.841 |
| Meta-epidemiology (narrow) | 0.001 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.004 |
| Bibliometrics | 0.012 | 0.012 |
| Science and technology studies | 0.013 | 0.106 |
| Scholarly communication | 0.053 | 0.106 |
| Open science | 0.011 | 0.031 |
| Research integrity | 0.043 | 0.045 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".