Guidance on review type selection for health technology assessments: key factors and considerations for deciding when to conduct a de novo systematic review, an update of a systematic review, or an overview of systematic reviews
Bibliographic record
Abstract
BACKGROUND: A systematic review (SR) helps us make sense of a body of research while minimizing bias and is routinely conducted to evaluate intervention effects in a health technology assessment (HTA). In addition to the traditional de novo SR, which combines the results of multiple primary studies, there are alternative review types that use systematic methods and leverage existing SRs, namely updates of SRs and overviews of SRs. This paper shares guidance that can be used to select the most appropriate review type to conduct when evaluating intervention effects in an HTA, with a goal to leverage existing SRs and reduce research waste where possible. PROCESS: We identified key factors and considerations that can inform the process of deciding to conduct one review type over the others to answer a research question and organized them into guidance comprising a summary and a corresponding flowchart. This work consisted of three steps. First, a guidance document was drafted by methodologists from two Canadian HTA agencies based on their experience. Next, the draft guidance was supplemented with a literature review. Lastly, broader feedback from HTA researchers across Canada was sought and incorporated into the final guidance. INSIGHTS: Nine key factors and six considerations were identified to help reviewers select the most appropriate review type to conduct. These fell into one of two categories: the evidentiary needs of the planned review (i.e., to understand the scope, objective, and analytic approach required for the review) and the state of the existing literature (i.e., to know the available literature in terms of its relevance, quality, comprehensiveness, currency, and findings). The accompanying flowchart, which can be used as a decision tool, demonstrates the interdependency between many of the key factors and considerations and aims to balance the potential benefits and challenges of leveraging existing SRs instead of primary study reports. CONCLUSIONS: Selecting the most appropriate review type to conduct when evaluating intervention effects in an HTA requires a myriad of factors to be considered. We hope this guidance adds clarity to the many competing considerations when deciding which review type to conduct and facilitates that decision-making process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.154 | 0.215 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.042 | 0.002 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".