Managing the Infodemic: Leveraging Deep Learning to Evaluate the Maturity Level of AI-Based COVID-19 Publications for Knowledge Surveillance and Decision Support
Bibliographic record
Abstract
ABSTRACT COVID-19 pandemic has taught us many lessons, including the need to manage the exponential growth of knowledge, fast-paced development or modification of existing AI models, limited opportunities to conduct extensive validation studies, the need to understand bias and mitigate it, and lastly, implementation challenges related to AI in healthcare. While the nature of the dynamic pandemic, resource limitations, and evolving pathogens were key to some of the failures of AI to help manage the disease, the infodemic during the pandemic could be a key opportunity that we could manage better. We share our research related to the use of deep learning methods to quantitatively and qualitatively evaluate AI-based COVID-19 publications which provides a unique approach to identify “mature” publications using a validated model and how that can be leveraged further by focused human-in-loop analysis. The study utilized research articles in English that were human-based, extracted from PubMed spanning the years 2020 to 2022. The findings highlight notable patterns in publication maturity over the years, with consistent and significant contributions from China and the United States. The analysis also emphasizes the prevalence of image datasets and variations in employed AI model types. To manage an infodemic during a pandemic, we provide a specific knowledge surveillance method to identify key scientific publications in near real-time. We hope this will enable data-driven and evidence-based decisions that clinicians, data scientists, researchers, policymakers, and public health officials need to make with time sensitivity while keeping humans in the loop.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.020 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".