Evaluating Trends in COVID-19 Research Activity in Early 2020: The Creation and Utilization of a Novel Open-Access Database
Bibliographic record
Abstract
Introduction The coronavirus disease 2019 (COVID-19) pandemic has been unprecedented in recent history. The rapid global spread has demonstrated how the emergence of a novel pathogen necessitates new information to advise both healthcare systems and policy-makers. The directives for the management of COVID-19 have been limited to infection control measures and treatment of patients, which has left physicians and researchers alone to navigate the massive amount of research being published while searching for evidence-based strategies to care for patients. To tackle this barrier, we launched CovidReview.ca, an open-access, continually updated, online platform that screens available COVID-19 research to determine higher quality publications. This paper uses data from this review process to explore the activity and trends of COVID-19 research worldwide over time, while specifically looking at the types of studies being published. Materials and Methods The literature search was conducted on PubMed. Search terms included "COVID-19", "severe acute respiratory syndrome coronavirus 2", "coronavirus 19", "SARS-COV-2", and "2019-nCoV". All articles captured by this strategy were reviewed by a minimum of two reviewers and categorized by type of research, relevant medical specialties, and type of publication. Criteria were developed to allow for inclusion or exclusion to the website. Due to the volume of research, only a level 1 (title and abstract) screen was performed. Results The time period for the analysis was January 17, 2020, to May 10, 2020. The total number of papers captured by the search criteria was 10,685, of which 2,742 were included on the website and 7,943 were excluded. The greatest increase in the types of studies over the 16 weeks was narrative review/expert opinion papers followed by case series/reports. Meta-analyses, systematic reviews, and randomized controlled trials remained the least published types of studies. Conclusions The surge of research that accompanied the COVID-19 pandemic is unparalleled in recent years. From our analysis, it is clear that case reports and narrative reviews were the most widely published, particularly in the earlier days of this pandemic. Continued research that falls higher on the evidence pyramid and is more applicable to clinical settings is warranted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.113 | 0.318 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.006 |
| Bibliometrics | 0.101 | 0.083 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.014 | 0.013 |
| Open science | 0.004 | 0.007 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.013 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".