Is convalescent plasma futile in COVID-19? A Bayesian re-analysis of the RECOVERY randomized controlled trial
Bibliographic record
Abstract
BACKGROUND: Randomized trials are generally performed from a frequentist perspective, which can conflate absence of evidence with evidence of absence. The RECOVERY trial evaluated convalescent plasma for patients hospitalized with coronavirus disease 2019 (COVID-19) and concluded that there was no evidence of an effect. Re-analysis from a Bayesian perspective is warranted. METHODS: Outcome data were extracted from the RECOVERY trial by serostatus and time of presentation. A Bayesian re-analysis with a wide variety of priors (vague, optimistic, sceptical, and pessimistic) was performed, calculating the posterior probability for: any benefit, an absolute risk difference of 0.5% (small benefit, number needed to treat 200), and an absolute risk difference of one percentage point (modest benefit, number needed to treat 100). RESULTS: Across all patients, when analysed with a vague prior, the likelihood of any benefit or a modest benefit with convalescent plasma was estimated to be 64% and 18%, respectively. The estimated chance of any benefit was 95% if presenting within 7 days of symptoms, or 17% if presenting after this. In patients without a detectable antibody response at presentation, the chance of any benefit was 85%. However, it was only 20% in patients with a detectable antibody response at presentation. CONCLUSIONS: Bayesian re-analysis suggests that convalescent plasma reduces mortality by at least one percentage point among the 39% of patients who present within 7 days of symptoms, and that there is a 67% chance of the same mortality reduction in the 38% who are seronegative at the time of presentation. This is in contrast to the results in people who already have antibodies when they present. This biologically plausible finding bears witness to the advantage of Bayesian analyses over misuse of hypothesis tests to inform decisions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".