Methodological Rigor in COVID-19 Clinical Research – A Systematic Review and Case-Control Analysis
Bibliographic record
Abstract
Abstract Objective To systematically evaluate the quality of reporting of currently available COVID-19 studies compared to historical controls. Design A systematic review and case-control analysis Data sources MEDLINE, Embase, and Cochrane Central Register of Controlled Trials until May 14, 2020 Study selection All original clinical literature evaluating COVID-19 or SARS-CoV2 were identified and 1:1 historical control of the same study type in the same published journal was matched from the previous year Data extraction Two independent reviewers screened titles, abstracts, and full-texts and independently assessed methodological quality using Cochrane Risk of Bias Tool, Newcastle- Ottawa Scale, QUADAS-2 Score, or case series checklist. Results 9895 titles and abstracts were screened and 686 COVID-19 articles were included in the final analysis in which 380 (55.4%) were case series, 199 (29.0%) were cohort, 63 (9.2%) were diagnostic, 38 (5.5%) were case-control, and 6 (0.9%) were randomized controlled trials. Overall, high quality/low-bias studies represented less than half of COVID-19 articles - 49.0% of case series, 43.9% of cohort, 31.6% of case-control, and 6.4% of diagnostic studies. We matched 539 control articles to COVID-19 articles from the same journal in the previous year for a final analysis of 1078 articles. The median time to acceptance was 13.0 (IQR, 5.0-25.0) days in COVID-19 articles vs. 110.0 (IQR, 71.0-156.0) days in control articles (p<0.0001). Overall, methodological quality was lower in COVID-19 articles with 220 COVID-19 articles of high quality (41.0%) vs. 392 control articles (73.3%, p<0.0001) with similar results when stratified by study design. In both unadjusted and adjusted logistic regression, COVID-19 articles were associated with lower methodological quality (odds ratio, 0.25; 95% CI, 0.20 to 0.33, p<0.0001). Conclusion Currently published COVID-19 studies were accepted more quickly and were found to be of lower methodological quality than comparative studies published in the same journal. Given the implications of these studies to medical decision making and government policy, greater effort to appropriately weigh the existing evidence in the context of emerging high-quality research is needed. Study registration PROSPERO: CRD42020187318
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.499 | 0.773 |
| Meta-epidemiology (narrow) | 0.004 | 0.004 |
| Meta-epidemiology (broad) | 0.020 | 0.018 |
| Bibliometrics | 0.038 | 0.026 |
| Science and technology studies | 0.004 | 0.007 |
| Scholarly communication | 0.016 | 0.010 |
| Open science | 0.007 | 0.007 |
| Research integrity | 0.007 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".