Development of a STandard reporting guideline for Evidence briefs for Policy (STEP): context and study protocol
Bibliographic record
Abstract
BACKGROUND: Evidence briefs for policy (EBP) draw on best-available data and research evidence (e.g., systematic reviews) to help clarify policy problems, frame options for addressing them, and identify implementation considerations for policymakers in a given context. An increasing number of governments, non-governmental organizations and research groups have been developing EBP on a wide variety of topics. However, the reporting characteristics of EBP vary across organizations due to a lack of internationally accepted standard reporting guidelines. This project aims to develop a STandard reporting guideline of Evidence briefs for Policy (STEP), which will encompass a reporting checklist and a STEP statement and a user manual. METHODS: We will refer to and adapt the methods recommended by the EQUATOR (Enhancing the QUAlity and Transparency Of health Research) network. The key actions include: (1) developing a protocol; (2) establishing an international multidisciplinary STEP working group (consisting of a Coordination Team and a Delphi Panel); (3) generating an initial draft of the potential items for the STEP reporting checklist through a comprehensive review of EBP-related literature and documents; (4) conducting a modified Delphi process to select and refine the reporting checklist; (5) using the STEP to evaluate published policy briefs in different countries; (6) finalizing the checklist; (7) developing the STEP statement and the user manual (8) translating the STEP into different languages; and (9) testing the reliability through real world use. DISCUSSION: Our protocol describes the development process for STEP. It will directly address what and how information should be reported in EBP and contribute to improving their quality. The decision-makers, researchers, journal editors, evaluators, and other stakeholders who support evidence-informed policymaking through the use of mechanisms like EBP will benefit from the STEP. Registration We registered the protocol on the EQUATOR network. ( https://www.equator-network.org/library/reporting-guidelines-under-development/#84 ).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Reporting · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
| gpt | Metaresearch Domain: Reporting · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.475 | 0.621 |
| Meta-epidemiology (narrow) | 0.005 | 0.008 |
| Meta-epidemiology (broad) | 0.007 | 0.013 |
| Bibliometrics | 0.020 | 0.017 |
| Science and technology studies | 0.007 | 0.008 |
| Scholarly communication | 0.014 | 0.015 |
| Open science | 0.011 | 0.014 |
| Research integrity | 0.015 | 0.022 |
| Insufficient payload (model declined to judge) | 0.031 | 0.020 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".