Development of a PRISMA extension for systematic reviews of health economic evaluations (PRISMA-EconEval): a project protocol
Bibliographic record
Abstract
BACKGROUND: Systematic reviews of health economic evaluations are key for evidence-based decisions but lack standardised reporting. This project aims to develop a Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) extension for systematic reviews of health economic evaluations (PRISMA-EconEval). METHODS: Project stages include the following: (1) scoping review, (2) Delphi surveys, (3) consensus meeting, (4) piloting, and (5) finalisation and dissemination. The project is overseen by the international multidisciplinary PRISMA-EconEval Management Group (PMG), Advisory Group, and Patient and Public Involvement Group. (1) The scoping review aims to identify candidate reporting items, with the protocol published elsewhere. The global applicability of these items to systematic reviews of health economic evaluations will be evaluated using sample papers from the scoping review, supplemented by nominations from the health economics community or other sources, where necessary. (2) A multi-round online Delphi survey will be conducted to achieve consensus on items for inclusion. A purposive sample of panellists (approximately 200) will be selected, ensuring representation of the following: health economists, systematic reviewers, information specialists, guideline developers, journal editors, healthcare decision-makers, research funders, and public representatives. Across two to three rounds, panellists will use a 1-9 scale to rate each candidate item's ability to represent the minimum required for reporting, be relevant to all systematic reviews of health economic evaluations, facilitate complete and transparent reporting, and support the quality assessment of both the review and included studies. (3) An online consensus meeting (approximately 30 participants) will refine the wording of items and resolve any disagreements by vote. (4) Health economists independent of the project will apply the draft guidelines to a sample of published studies and identify practical challenges. (5) The PMG will meet to finalise the wording and presentation of the reporting items, ensure consistency with PRISMA 2020, and produce an explanation and elaboration document. Dissemination channels will include peer-reviewed health economics journals, conferences, and the EQUATOR network. DISCUSSION: PRISMA-EconEval aims to improve clarity, consistency, transparency, quality, and overall value of systematic reviews of health economic evaluations. This will benefit researchers, peer reviewers, editors, decision-makers, and ultimately patients and the public through supporting decisions on healthcare resource allocation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.175 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.011 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".