Consistency and Completeness of Retractions in Public Health Research on COVID-19
Bibliographic record
Abstract
Caitlin J. Bakker,<sup>1,2</sup> Erin E. Reardon,<sup>3</sup> Sarah Jane Brown,<sup>4</sup> Nicole Theis-Mahon,<sup>4</sup> Sara Schroter,<sup>5,6</sup> Lex Bouter,<sup>7,8</sup> Maurice P. Zeegers<sup>2</sup> <h4>Objective </h4> The scientific community responded rapidly to COVID-19, producing over 200,000 publications in 1 year.<sup>1</sup> This speed brought challenges, including a higher retraction rate.<sup>2</sup> While retraction helps correct the scientific record, retracted status may be inconsistently presented and notices may be incomplete.<sup>3</sup> During a pandemic, inconsistency and incompleteness can have immediate and long-term health impacts. We evaluated retraction presentation consistency and notice completeness for COVID-19 vs non–COVID-19 publications. We describe reasons for retraction and time from publication to retraction. <h4>Design </h4> In March 2023, we retrieved retracted publications categorized as research articles or clinical studies in the subject area of public health and safety from Retraction Watch. A previous study focused on all retracted publications in this period<sup>3</sup>; this is a sub-study comparing COVID-19 with non–COVID-19 publications. Between April 28 and June 6, 2023, we assessed consistency in 11 databases (Academia.edu, CINAHL, Embase.com, Ovid Embase, Ovid Medline, PubMed, ResearchGate, SciHub, Scopus, Web of Science, and publisher websites) using 12 criteria from the International Committee of Medical Journal Editors and National Library of Medicine and notice completeness using 17 criteria from Retraction Watch and the Committee on Publication Ethics. Criteria were scored 0 if missing or 1 if present; partial scores were assigned when evaluating multicomponent criteria, such as bidirectional links. To ensure consistent scoring, a random subset of 21% (92 of 441) of retracted publications were independently reviewed by 2 researchers. Each researcher extracted data using Qualtrics forms, and scoring discrepancies were resolved by consensus. Following this calibration phase, remaining publications were extracted by a single reviewer. Kruskal-Wallis tests assessed differences in scores between COVID-19 and non–COVID-19 publications. <h4>Results</h4> Of 441 publications, 47 were about COVID-19 and 394 were not. COVID-19 publications were published between 2019 and 2022, while non-COVID 19 publications were published between 1978 and 2022. COVID-19 publications were most frequently retracted due to concerns about reliability of data or results (15 [31.9%]) compared with plagiarism (82 [20.8%]) for non–COVID-19 publications. The median time between publication and retraction was 120 (IQR, 15-196) days for COVID-19 publications and 326 (IQR, 124-789) days for non–COVID-19 publications (<i>P</i> < .001). Across 11 databases, 41.2% (110 of 267) of records retrieved for retracted COVID-19 publications were marked as retracted compared with 47.6% (1225 of 2574) for non–COVID-19 publications. There was no statistically significant difference between consistency or completeness scores for COVID-19 vs non–COVID-19 retracted publications (<b>Table 25-1062</b>). No publications met all criteria. https://assets.underline.io/markdown_image/1/image/edc419e41ec09f04154b15739dea0699.png <h4>Conclusions </h4> Incomplete and inconsistent information poses challenges for researchers and practitioners, undermining trust in scientific literature. We found no association between publications being about COVID-19 and the consistency or completeness of retraction information; however, publications about COVID-19 appeared to be retracted more quickly. <h4>References</h4> 1. Shimray SR. Research done wrong: a comprehensive investigation of retracted publications in COVID-19. <i>Account Res</i>. 2022;30(7):393-406. doi:10.1080/08989621.2021.2014327 2. Yeo-Teh NSL, Tang BL. An alarming retraction rate for scientific publications on coronavirus disease 2019 (COVID-19). <i>Account Res</i>. 2020;28(1):47-53. doi:10.1080/08989621.2020.1782203 3. Bakker CJ, Reardon EE, Brown SJ, et al. Identification of retracted publications and completeness of retraction notices in public health. <i>J Clin Epidemiol</i>. 2024;173:111427. doi:10.1016/j.jclinepi.2024.111427 <sup>1</sup>University of Regina, Regina, Saskatchewan, Canada, caitlin.bakker@uregina.ca; <sup>2</sup>Maastricht University, Maastricht, the Netherlands; <sup>3</sup>Emory University, Atlanta, GA, US; <sup>4</sup>University of Minnesota, Minneapolis, MN, US; <sup>5</sup><i>BMJ</i>, London, UK; <sup>6</sup>London School of Hygiene and Tropical Medicine, London, UK; <sup>7</sup>Amsterdam University Medical Center, Amsterdam, the Netherlands; <sup>8</sup>Vrije Universiteit Amsterdam, Amsterdam, the Netherlands. <h4>Conflict of Interest Disclosures</h4> Caitlin J. Bakker is cochair of the National Information Standards Organization Communication of Retractions, Removals and Expressions of Concern Standing Committee. Lex Bouter is a member of the Peer Review Congress Advisory Board but was not involved in the review or decision for this abstract. No other disclosures were reported. <h4>Funding/Support</h4> This research is part of an ongoing PhD collaboration between The BMJ (British Medical Journal) and the team Meta-Research at Maastricht University (UM) on the responsible conduct of publishing scientific research. The BMJ is published by BMJ Group, a wholly owned subsidiary of the British Medical Association. UM is a public legal entity in the Netherlands. This study is part of Caitlin Bakker’s self-funded BMJ/UM PhD. No exchange of funds has taken place for this research project. <h4>Role of the Funder/Sponsor</h4> The authors are wholly responsible for the design and conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the abstract; and decision to submit the abstract for presentation. <h4>Disclaimer</h4> All authors express their own opinions and not necessarily that of their employers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.022 | 0.027 |
| Science and technology studies | 0.003 | 0.030 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".