MétaCan
Menu
Back to cohort
Record W3128992408 · doi:10.1371/journal.pone.0246427

General medical publications during COVID-19 show increased dissemination despite lower validation

2021· article· en· W3128992408 on OpenAlexafffund
Nan Gai, Kazuyoshi Aoyama, David Faraoni, Neil M. Goldenberg, David Levin, Jason T. Maynes, Mark J. McVey, Farrukh Munshey, Asad Ali Siddiqui, Timothy Switzer, Benjamin E. Steinberg

Bibliographic record

VenuePLoS ONE · 2021
Typearticle
Languageen
FieldDecision Sciences
TopicAcademic Publishing and Open Access
Canadian institutionsSickKids FoundationInstitute for Clinical Evaluative SciencesMental Health Research CanadaHospital for Sick Children
FundersHospital for Sick Children
KeywordsCoronavirus disease 2019 (COVID-19)PandemicObservational studyBibliometricsSevere acute respiratory syndrome coronavirus 2 (SARS-CoV-2)2019-20 coronavirus outbreakMEDLINEMedicineLibrary scienceComputer sciencePolitical sciencePathologyOutbreak

Abstract

fetched live from OpenAlex

BACKGROUND: The COVID-19 pandemic has yielded an unprecedented quantity of new publications, contributing to an overwhelming quantity of information and leading to the rapid dissemination of less stringently validated information. Yet, a formal analysis of how the medical literature has changed during the pandemic is lacking. In this analysis, we aimed to quantify how scientific publications changed at the outset of the COVID-19 pandemic. METHODS: We performed a cross-sectional bibliometric study of published studies in four high-impact medical journals to identify differences in the characteristics of COVID-19 related publications compared to non-pandemic studies. Original investigations related to SARS-CoV-2 and COVID-19 published in March and April 2020 were identified and compared to non-COVID-19 research publications over the same two-month period in 2019 and 2020. Extracted data included publication characteristics, study characteristics, author characteristics, and impact metrics. Our primary measure was principal component analysis (PCA) of publication characteristics and impact metrics across groups. RESULTS: We identified 402 publications that met inclusion criteria: 76 were related to COVID-19; 154 and 172 were non-COVID publications over the same period in 2020 and 2019, respectively. PCA utilizing the collected bibliometric data revealed segregation of the COVID-19 literature subset from both groups of non-COVID literature (2019 and 2020). COVID-19 publications were more likely to describe prospective observational (31.6%) or case series (41.8%) studies without industry funding as compared with non-COVID articles, which were represented primarily by randomized controlled trials (32.5% and 36.6% in the non-COVID literature from 2020 and 2019, respectively). CONCLUSIONS: In this cross-sectional study of publications in four general medical journals, COVID-related articles were significantly different from non-COVID articles based on article characteristics and impact metrics. COVID-related studies were generally shorter articles reporting observational studies with less literature cited and fewer study sites, suggestive of more limited scientific support. They nevertheless had much higher dissemination.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.024
metaresearch head score (Gemma)0.133
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Bibliometrics
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.976
Threshold uncertainty score0.125

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0240.133
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0420.080
Science and technology studies0.0010.002
Scholarly communication0.0080.005
Open science0.0010.004
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.134
GPT teacher head0.401
Teacher spread0.267 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations13
Published2021
Admission routes2
Has abstractyes

Explore more

Same venuePLoS ONESame topicAcademic Publishing and Open AccessFrench-language works237,207