Methodological and reporting quality of systematic and rapid reviews on human mpox and their utility during a public health emergency
Bibliographic record
Abstract
Introduction: Evidence syntheses were rapidly produced during the 2022 mpox outbreak despite a lack of studies. The aim of this methodological study was to assess the quality and utility of the evidence syntheses produced during the first 6 months of the outbreak compared to those published before it. Methods: Human mpox evidence syntheses available before December 31, 2022 were retrieved from PubMed, Scopus, EuropePMC, SSRN, and arXiv. Study characteristics, utility, methodological, and reporting quality (AMSTAR-2 and PRISMA) were contrasted between syntheses produced before the 2022 outbreak (historical) and during the first 6 months (new). Results were synthesized narratively. Results: Twenty-six evidence syntheses were included; two historical systematic reviews (SRs) and 24 new SRs, rapid reviews, scoping reviews, and mislabelled syntheses. Median time from search to publication/preprint post date was 68 and 6 weeks for historical and new syntheses, respectively. Among the new syntheses, 8% (2/24) did not include evidence from the 2022 outbreak, 33% (8/24) included only new evidence and 58% (14/24) included both new and historical evidence. Only 29% of new syntheses contrasted findings between new and historical evidence. Methodological quality was critically low for 100% of historical syntheses and 92% of new syntheses and the remainder (8%) were low. Reporting quality was poor with a median of 10.5 (range 10-11) and 11.5 (range 4-21) of 27 items reported sufficiently by historical and new syntheses, respectively. Conclusions: Evidence syntheses take time to produce and during an emergent outbreak they are often outdated at the time of publication and suffer from poor adherence to methodological and reporting guidelines. Overlapping content and few new studies resulted in minimal added value to the mpox literature. Strategies to reduce duplication and mechanisms to produce and disseminate continuously updated living evidence syntheses need to be explored to support decision-makers responding to an emergency.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.057 | 0.060 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".