Open science interventions proposed or implemented to assess researcher impact: a scoping review
Bibliographic record
Abstract
Background Several open science-promoting initiatives have been proposed to improve the quality of biomedical research, including initiatives for assessing researchers’ open science behaviour as criteria for promotion or tenure. Yet there is limited evidence to judge whether the interventions are effective. This review aimed to summarise the literature, identifying open science practices related to researcher assessment, and map the extent of evidence of existing interventions implemented to assess researchers and research impact. Methods A scoping review using the Joanna Briggs Institute Scoping Review Methodology was conducted. We included all study types that described any open science practice-promoting initiatives proposed or implemented to assess researchers and research impact, in health sciences, biomedicine, psychology, and economics. Data synthesis was quantitative and descriptive. Results Among 18,020 identified documents, 27 articles were selectedfor analysis. Most of the publications were in the field of health sciences (n = 10), and were indicated as research culture, perspective, commentary, essay, proceedings of a workshop, research article, world view, opinion, research note, editorial, report, and research policy articles (n = 22). The majority of studies proposed recommendations to address problems regarding threats to research rigour and reproducibility that were multi-modal (n = 20), targeting several open science practices. Some of the studies based their proposed recommendations on further evaluation or extension of previous initiatives. Most of the articles (n = 20) did not discuss implementation of their proposed intervention. Of the 27 included articles, 10 were cited in policy documents, with The Leiden Manifesto being the most cited (104 citations). Conclusion This review provides an overview of proposals to integrate open science into researcher assessment. The more promising ones need evaluation and, where appropriate, implementation. Study registration https://osf.io/ty9m7
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchOpen science Domain: Incentives · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
| gpt | MetaresearchOpen science Domain: Evaluation · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.175 | 0.443 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.005 | 0.009 |
| Bibliometrics | 0.030 | 0.030 |
| Science and technology studies | 0.004 | 0.004 |
| Scholarly communication | 0.011 | 0.014 |
| Open science | 0.005 | 0.009 |
| Research integrity | 0.006 | 0.004 |
| Insufficient payload (model declined to judge) | 0.011 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".