Curated Collections for Educators: Five Key Papers on Evaluating Digital Scholarship
Bibliographic record
Abstract
Traditionally, scholarship that was recognized for promotion and tenure consisted of clinical research, bench research, and grant funding. Recent trends have allowed for differing approaches to scholarship, including digital publication. As increasing numbers of trainees and faculty turn to online educational resources, it is imperative to critically evaluate these resources. This article summarizes five key papers that address the appraisal of digital scholarship and describes their relevance to junior clinician educators and faculty developers. In May 2017, the Academic Life in Emergency Medicine Faculty Incubator program focused on the topic of digital scholarship, providing and discussing papers relevant to the topic. We augmented this list of papers with further suggestions by guest experts and by an open call via Twitter for other important papers. Through this process, we created a list of 38 papers in total on the topic of evaluating digital scholarship. In order to determine which of these papers best describe how to evaluate digital scholarship, the authorship group assessed the papers using a modified Delphi approach to build consensus. In this paper we present the five most highly rated papers from our process about evaluating digital scholarship. We summarize each paper and discuss its specific relevance to junior faculty members and to faculty developers. These papers provide a framework for assessing the quality of digital scholarship, so that junior faculty can recommend high-quality educational resources to their trainees. These papers help guide educators on how to produce high quality digital scholarship and maximize recognition and credit in respect to receiving promotion and tenure.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.060 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".