Can AI learn to identify systematic reviews on the effectiveness of public health interventions?
Bibliographic record
Abstract
Abstract Issue/problem Health Evidence™ aims to make it easier for public health professionals and decision-makers to use evidence in their programs and policies. We provide access to over 7,000 critically appraised systematic reviews on the effectiveness of public health interventions. On average, 8,000-10,000 records are screened each month to identify around 50 relevant reviews that are critically appraised and uploaded to the registry. As the number of published reviews continues to grow each year, maintaining an up-to-date Registry is becoming increasingly resource intensive. Description of the problem Artificial intelligence (AI) may be one way to ensure maintenance of this Registry continues to be feasible. It is important that the use of AI is accurate and efficient to support monthly relevance screening for the Health Evidence ™ Registry. To assess if the use of AI is appropriate in this context, the team uploaded a large, labelled training set (n = 43,273) of relevant and non-relevant records to DistillerSR and used the AI Preview and Rank function to predict the probability of relevance to the Registry. We identified an optimal threshold (0.17) that automatically removes the greatest number of records with minimal classification errors. We then tested this threshold on one year of manually screened records (n = 89,832) from May 2019 to June 2020 to ensure continued accuracy. Results Using AI to support our monthly relevance screening has reduced our monthly manual screening burden by over 77% (from ∼9,000 references manually screened per month to ∼2,000) saving upwards of 10 hours staff time per month! Lessons The use of AI shows promise to help improve the feasibility of maintaining a large Registry of quality appraised synthesis level evidence relevant for public health decision making. Key messages Health Evidence™ provides easy access to high quality synthesis evidence relevant for public health. The use of AI shows promise for helping to automate the Health Evidence™ monthly update and improve the feasibility of maintaining a public health registry of quality appraised synthesis evidence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.213 | 0.744 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.006 |
| Bibliometrics | 0.028 | 0.016 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.011 | 0.011 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".