Methods for Developing Evidence Reviews in Short Periods of Time: A Scoping Review
Bibliographic record
Abstract
INTRODUCTION: Rapid reviews (RR), using abbreviated systematic review (SR) methods, are becoming more popular among decision-makers. This World Health Organization commissioned study sought to summarize RR methods, identify differences, and highlight potential biases between RR and SR. METHODS: Review of RR methods (Key Question 1 [KQ1]), meta-epidemiologic studies comparing reliability/ validity of RR and SR methods (KQ2), and their potential associated biases (KQ3). We searched Medline, EMBASE, Cochrane Library, grey literature, and checked reference lists, used personal contacts, and crowdsourcing (e.g. email listservs). Selection and data extraction was conducted by one reviewer (KQ1) or two reviewers independently (KQ2-3). RESULTS: Across all KQs, we identified 42,743 citations through the literature searches. KQ1: RR methods from 29 organizations were reviewed. There was no consensus on which aspects of the SR process to abbreviate. KQ2: Studies comparing the conclusions of RR and SR (n = 9) found them to be generally similar. Where major differences were identified, it was attributed to the inclusion of evidence from different sources (e.g. searching different databases or including different study designs). KQ3: Potential biases introduced into the review process were well-identified although not necessarily supported by empirical evidence, and focused mainly on selective outcome reporting and publication biases. CONCLUSION: RR approaches are context and organization specific. Existing comparative evidence has found similar conclusions derived from RR and SR, but there is a lack of evidence comparing the potential of bias in both evidence synthesis approaches. Further research and decision aids are needed to help decision makers and reviewers balance the benefits of providing timely evidence with the potential for biased findings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Methods · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | low |
| gpt | Metaresearch Domain: Methods · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.369 | 0.331 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.045 | 0.007 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.004 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".