Using artificial intelligence to support and streamline rapid systematic evidence reviews
Bibliographic record
Abstract
Abstract Issue The Rapid Evidence Service, initiated by the National Collaborating Centre for Methods and Tools (NCCMT) during COVID-19, supports public health decision making by conducting rapid reviews on priority topics. Description of issue Integral to rapid reviews is an expedited timeline but the quantity of available literature for most public health review questions takes significant time to screen manually. NCCMT integrated 4 artificial intelligence (AI) features into the screening process. DAISY Rank applies predictions learned from manual screening patterns to re-order remaining studies, with most relevant appearing first. AI Screening automatically screens remaining studies based on prediction scores. Check for Screening Errors and Re-Rank Report use previous screening patterns to identify studies that were potentially falsely excluded and predict the total number of included studies, respectively. These features were tested by comparing results provided by AI with those produced manually for select test sets. Results NCCMT used AI to support and expedite screening, assess screening progress, and/or minimize risk of inappropriately excluding studies for 35 rapid reviews on 20 topics. Using DAISY Rank enabled one screener to review over 4000 references in 9 hours, compared to a different review, where the same amount of screening took 28 hours without DAISY Rank. AI Screening correctly excluded up to 80% of irrelevant search results across reviews. Check for Screening Errors identified 37 potential includes manually excluded in one review; these were reviewed and 3 were included. Re-Rank Report allowed NCCMT to re-allocate staff to subsequent steps in the review process when most included studies were identified. Lessons Integrating AI features into screening led to less time required, better anticipated timelines, more accurate staff allocation and reduced errors. More rigorous study of AI best practices is needed to continue to improve rapid review method efficiencies. Key messages • Rapid reviews can be an important source of evidence for decision makers if they can be completed quickly but maintain rigor and accuracy. • AI holds promise as a way to improve screening efficiency.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.278 | 0.648 |
| Meta-epidemiology (narrow) | 0.003 | 0.004 |
| Meta-epidemiology (broad) | 0.005 | 0.005 |
| Bibliometrics | 0.026 | 0.015 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.014 | 0.010 |
| Open science | 0.005 | 0.008 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".