Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation
Bibliographic record
Abstract
ABSTRACT Objective To compare results reporting and the presence of spin in COVID-19 study preprints with their finalized journal publications Design Cross-sectional Setting International medical literature Participants Preprints and final journal publications of 67 interventional and observational studies of COVID-19 treatment or prevention from the Cochrane COVID-19 Study Register published between March 1, 2020 and October 30, 2020 Main outcome measures Study characteristics and discrepancies in 1) Results reporting (number of outcomes, outcome descriptor, measure (e.g., PCR test), metric (e.g., mean change from baseline), assessment time point (e.g., 1 week post treatment), data reported (e.g., effect estimate and measures of precision), reported statistical significance of result, type of statistical analysis (e.g., chi-squared test), subgroup analyses (if any), whether outcome was identified as primary or secondary and 2) Spin (reporting practices that distort the interpretation of results so that results are viewed more favorably). Results Of 67 included studies, 23 (34%) had no discrepancies in results reporting between preprints and journal publications. Fifteen (22%) studies had at least one outcome that was included in the journal publication, but not the preprint; 8 (12%) had at least one outcome that was reported in the preprint only. For outcomes that were reported in both preprints and journals, common discrepancies were differences in numerical values and statistical significance, additional statistical tests and subgroup analyses conducted in journal publications, and longer follow-up times for outcome assessment in journal publications. At least one instance of spin occurred in both preprints and journals in 23 / 67 (34%) studies, the preprint only in 5 (7%) studies, and the journal publications only in 2 (3%) of studies. Spin was removed between the preprint and journal publication in 5/67 (7%) studies; but added in 1/67 (1%) study. Conclusions The COVID-19 preprints and their subsequent journal publications were largely similar in reporting of study characteristics, outcomes and spin. All COVID-19 studies published as preprints and journal publications should be critically evaluated for discrepancies and spin. EQUATOR REPORTING GUIDELINE STROBE What is already known on this topic Selective and incomplete reporting of results and spin are threats to the trustworthiness and validity of research. These reporting practices could be particularly dangerous for users of COVID-19 research as they can inflate the efficacy of interventions and underestimate harms. Given the high prevalence, visibility, and potentially rapid implementation of COVID-19 research published as preprints, it is important to compare components of results reporting and the presence of spin in COVID-19 studies on treatment or prevention that are published both as preprints and journal publications. What this study adds This comparison of 67 COVID-19 preprints related to treatment or prevention and their subsequent journal publications found they were largely similar in reporting of study characteristics, components of results reporting and spin in interpretation. Even a few important discrepancies could impact decision making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.302 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".