Errors in the Conduct of Systematic Reviews of Pharmacological Interventions for Irritable Bowel Syndrome
Bibliographic record
Abstract
OBJECTIVES: Systematic reviews and meta-analyses are integral to evidence-based clinical decision making. Although flawed systematic reviews could compromise optimal decision making, their accuracy has received limited investigation. We assessed conduct of systematic reviews of pharmaceutical interventions for irritable bowel syndrome (IBS). METHODS: We searched MEDLINE, EMBASE, and the Cochrane Database of Systematic Reviews (up to June 2008) to identify and replicate all published systematic reviews and meta-analyses that examined pharmacological interventions for IBS. We identified trials appropriately and inappropriately included according to the investigators' own eligibility criteria and eligible trials the investigators failed to include, and assessed the accuracy of dichotomous data extraction from all truly eligible trials. We conducted meta-analyses of accurate data from all truly eligible trials, and examined the differences between these accurate estimates and those reported by the authors. RESULTS: The search strategy identified 120 citations, and 13 appeared to be relevant. Five systematic reviews did not extract dichotomous data, leaving eight eligible for inclusion. In five of the eight meta-analyses 13-29% of included trials were ineligible according to investigators' criteria, constituting 8-26% of included patients. Six of the meta-analyses missed 17 separate published eligible trials; 3-11% of eligible patients were, as a result, not included. All eight meta-analyses contained errors in dichotomous data extraction, in 29-100% of truly eligible trials, leading to errors in 15 of 16 reported pooled treatment effects. There was a > or =10% relative difference in treatment effects between the reported and recalculated summary statistic in five (31%) cases, and a change in the statistical significance of the recalculated summary statistic in a further four (25%) cases. CONCLUSIONS: We found many errors in both application of eligibility criteria and dichotomous data extraction in the eight meta-analyses studied. Independent verification of systematic reviews and meta-analyses may be required for full confidence in their results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.704 | 0.908 |
| Meta-epidemiology (narrow) | 0.005 | 0.007 |
| Meta-epidemiology (broad) | 0.015 | 0.012 |
| Bibliometrics | 0.030 | 0.034 |
| Science and technology studies | 0.003 | 0.010 |
| Scholarly communication | 0.014 | 0.010 |
| Open science | 0.008 | 0.008 |
| Research integrity | 0.009 | 0.006 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".