MétaCan
Menu
Back to cohort
Record W7115741523 · doi:10.48448/sc28-0c69

Retraction of Systematic Reviews and Clinical Practice Guidelines

2025· other· W7115741523 on OpenAlexaboutno aff

Bibliographic record

VenueUnderline Science Inc. · 2025
Typeother
Language
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsClinical PracticeSystematic reviewMEDLINEDescriptive statisticsAlternative medicine

Abstract

fetched live from OpenAlex

Ivan D. Florez,<sup>1,2</sup> Alberto Henriquez,<sup>3,4</sup> Andrés F. Estupinan-Bohorquez<sup>3,5</sup> <h4>Objective </h4> We described the influence of retracted systematic reviews and meta-analyses (SRMAs) on clinical practice guidelines (CPGs) and the characteristics of retractions in CPGs. <h4>Design </h4> This cross-sectional study was conducted in 2 stages based on searches focused on the Retraction Watch (RW) database and MEDLINE from inception to November 30, 2024. In the first stage, we included SRMAs. We described the reasons for retractions, and we recategorized them based on our assessment. We identified the CPGs that cited the SRMA in the Google Scholar database. In the second stage, we included the retracted CPGs, described the reasons for retractions, and recategorized them into ethical or nonethical reasons based on our assessment. Nonethical reasons were categorized as editorial or administrative or outdated guidelines, while ethical reasons were reported according to RW categories. We used descriptive statistics to summarize the findings. <h4>Results </h4> In the first stage, we included 377 SRMAs, of which 211 (56.0%) were retracted due to peer review or publication manipulation (eg, detected “fake” reviewers); 30 (8.0%), due to duplicate or redundant publication; and 136 (36.1%), due to intellectual or authorship disputes, plagiarism, outdated publication, retraction of included studies, methodologic or data errors, and conflicts of interests. For 49 (13.0%), specific reasons were not provided. Of the retracted SRMAs, 41 (10.9%) were cited in CPGs; 19 (46.3%) of these SRMAs were retracted due to research integrity issues and 12 (29.3%), due to data errors or being outdated. For 10 (24.4%), specific reasons were not provided. Most retractions were due to manipulation of the publication or peer review process. The median time between publication and retraction of the SRMA used in CPGs was 12.0 (IQR, 3.5-25.0) months, and the median number of SRMA citations was 40 (IQR, 22-191). In the second stage, we included 36 CPGs of the 138 potential CPGs identified. Nine CPGs (25.0%) were retracted because of ethical reasons and 22 (61.1%) for nonethical reasons; the rest had no available information. The most common ethical reasons were plagiarism, authorship or intellectual property disputes, lack of disclosure of conflicts of interest, and discrepancies between the content and the cited evidence. Among the 22 CPGs retracted for nonethical reasons, 11 were due to dual publication or incorrect citations and 9 were due to outdated recommendations. The median publication to retraction time was 10 (range, 3-96) months. All of these CPGs were cited after their retraction date, and in all cases, the citations were used to support the background of the research studies. <h4>Conclusions </h4> Retracted SRMAs have been informing CPGs, which provide recommendations in practice and policy. The retraction of CPGs has been neglected. The RW database should be revised according to the specificities of CPGs. The most concerning reasons are ethical. Retracted CPGs continue to be cited after their retraction, mainly to inform the background sections of articles. <sup>1</sup>Department of Pediatrics, University of Antioquia, Medellín, Colombia, ivan.florez@udea.edu.co; <sup>2</sup>School of Rehabilitation Science, McMaster University, Hamilton, Ontario, Canada; <sup>3</sup>Universidad del Norte, Barranquilla, Colombia; <sup>4</sup>Universidad Metropolitana, Barranquilla, Colombia; <sup>5</sup>EPICLINICA SAS, Barranquilla, Colombia. <h4>Conflict of Interest Disclosures</h4> None reported.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearchResearch integrity
Domain: Evaluation · Genre: Other
About the Canadian research system: no · About a Canadian topic: no
Not applicablelow
gptResearch integrity
Domain: not available · Genre: Other
About the Canadian research system: no · About a Canadian topic: no
Other designlow
models splitAgreement compares identical category sets and study designs across arms.

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.088
metaresearch head score (Gemma)0.425
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Science and technology studies, Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Insufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.505
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0880.425
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0050.001
Bibliometrics0.0030.006
Science and technology studies0.0010.007
Scholarly communication0.0010.002
Open science0.0020.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.225
GPT teacher head0.510
Teacher spread0.285 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

MetaresearchResearch integrity

The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.

Study designNot applicable · Other design
DomainEvaluation
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueUnderline Science Inc.CategoryMetaresearchFrench-language works237,207