COVID-19 literature surveillance—A framework to manage the literature and support evidence-based decision-making on a rapidly evolving public health topic
Bibliographic record
Abstract
Background: The coronavirus disease 2019 (COVID-19) pandemic has led to a rapid surge of literature on severe acute respiratory syndrome coronavirus 2 and the wider impacts of the pandemic. Research on COVID-19 has been produced at an unprecedented rate, and the ability to stay on top of the most relevant evidence is top priority for clinicians, researchers, public health professionals and policymakers. This article presents a knowledge synthesis methodology developed and used by the Public Health Agency of Canada for managing and maintaining a literature surveillance system to identify, characterize, categorize and disseminate COVID-19 evidence daily. Methods: The Daily Scan of COVID-19 Literature project comprised a systematic process involving four main steps: literature search; screening for relevance; classification and summarization of studies; and disseminating a daily report. Results: As of the end of March 2022 there were approximately 300,000 COVID-19 and pandemic-related citations in the COVID-19 database, of which 50%-60% were primary research. Each day, a report of all new COVID-19 citations, literature highlights and a link to the updated database was generated and sent to a mailing list of over 200 recipients including federal, provincial and local public health agencies and academic institutions. Conclusion: This central repository of COVID-19 literature was maintained in real time to aid in accelerated evidence synthesis activities and support evidence-based decision-making during the pandemic response in Canada. This systematic process can be applied to future rapidly evolving public health topics that require the continuous evaluation and dissemination of evidence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.488 | 0.532 |
| Meta-epidemiology (narrow) | 0.004 | 0.005 |
| Meta-epidemiology (broad) | 0.009 | 0.009 |
| Bibliometrics | 0.157 | 0.100 |
| Science and technology studies | 0.014 | 0.014 |
| Scholarly communication | 0.047 | 0.033 |
| Open science | 0.016 | 0.035 |
| Research integrity | 0.009 | 0.009 |
| Insufficient payload (model declined to judge) | 0.007 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".