Retraction of Systematic Reviews and Clinical Practice Guidelines
Bibliographic record
Abstract
Ivan D. Florez,<sup>1,2</sup> Alberto Henriquez,<sup>3,4</sup> Andrés F. Estupinan-Bohorquez<sup>3,5</sup> <h4>Objective </h4> We described the influence of retracted systematic reviews and meta-analyses (SRMAs) on clinical practice guidelines (CPGs) and the characteristics of retractions in CPGs. <h4>Design </h4> This cross-sectional study was conducted in 2 stages based on searches focused on the Retraction Watch (RW) database and MEDLINE from inception to November 30, 2024. In the first stage, we included SRMAs. We described the reasons for retractions, and we recategorized them based on our assessment. We identified the CPGs that cited the SRMA in the Google Scholar database. In the second stage, we included the retracted CPGs, described the reasons for retractions, and recategorized them into ethical or nonethical reasons based on our assessment. Nonethical reasons were categorized as editorial or administrative or outdated guidelines, while ethical reasons were reported according to RW categories. We used descriptive statistics to summarize the findings. <h4>Results </h4> In the first stage, we included 377 SRMAs, of which 211 (56.0%) were retracted due to peer review or publication manipulation (eg, detected “fake” reviewers); 30 (8.0%), due to duplicate or redundant publication; and 136 (36.1%), due to intellectual or authorship disputes, plagiarism, outdated publication, retraction of included studies, methodologic or data errors, and conflicts of interests. For 49 (13.0%), specific reasons were not provided. Of the retracted SRMAs, 41 (10.9%) were cited in CPGs; 19 (46.3%) of these SRMAs were retracted due to research integrity issues and 12 (29.3%), due to data errors or being outdated. For 10 (24.4%), specific reasons were not provided. Most retractions were due to manipulation of the publication or peer review process. The median time between publication and retraction of the SRMA used in CPGs was 12.0 (IQR, 3.5-25.0) months, and the median number of SRMA citations was 40 (IQR, 22-191). In the second stage, we included 36 CPGs of the 138 potential CPGs identified. Nine CPGs (25.0%) were retracted because of ethical reasons and 22 (61.1%) for nonethical reasons; the rest had no available information. The most common ethical reasons were plagiarism, authorship or intellectual property disputes, lack of disclosure of conflicts of interest, and discrepancies between the content and the cited evidence. Among the 22 CPGs retracted for nonethical reasons, 11 were due to dual publication or incorrect citations and 9 were due to outdated recommendations. The median publication to retraction time was 10 (range, 3-96) months. All of these CPGs were cited after their retraction date, and in all cases, the citations were used to support the background of the research studies. <h4>Conclusions </h4> Retracted SRMAs have been informing CPGs, which provide recommendations in practice and policy. The retraction of CPGs has been neglected. The RW database should be revised according to the specificities of CPGs. The most concerning reasons are ethical. Retracted CPGs continue to be cited after their retraction, mainly to inform the background sections of articles. <sup>1</sup>Department of Pediatrics, University of Antioquia, Medellín, Colombia, ivan.florez@udea.edu.co; <sup>2</sup>School of Rehabilitation Science, McMaster University, Hamilton, Ontario, Canada; <sup>3</sup>Universidad del Norte, Barranquilla, Colombia; <sup>4</sup>Universidad Metropolitana, Barranquilla, Colombia; <sup>5</sup>EPICLINICA SAS, Barranquilla, Colombia. <h4>Conflict of Interest Disclosures</h4> None reported.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchResearch integrity Domain: Evaluation · Genre: Other About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | Research integrity Domain: not available · Genre: Other About the Canadian research system: no · About a Canadian topic: no | Other design | low |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.088 | 0.425 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.001 | 0.007 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".