MétaCan
Menu
Back to cohort
Record W4416807630 · doi:10.1093/humrep/deaf217

Safeguarding WHO guideline recommendations through strengthened scientific integrity to advance global health

2025· article· en· W4416807630 on OpenAlexaff
Gitau Mburu, Nancy Santesso, Romina Brignardello‐Petersen, Cindy Farquhar, Richard H. Kennedy, James Kiarie

Bibliographic record

VenueHuman Reproduction · 2025
Typearticle
Languageen
FieldMedicine
TopicPharmaceutical Quality and Counterfeiting
Canadian institutionsMcMaster UniversityImpact
FundersUNICEFWorld Health Organization
KeywordsSafeguardingGuidelineScientific integrityGlobal healthMEDLINEPublic health

Abstract

fetched live from OpenAlex

The release of the first WHO guidelines on prevention, diagnosis, and treatment of infertility (World Health Organization, 2025) is an important milestone in global public health, as infertility affects one in six people during their lifetime (World Health Organization, 2023). The guideline’s emphasis on evidence-based medicine is timely as the use of non-evidence-based interventions in medically assisted reproduction is common (van de Wiel et al., 2020). The guideline recommendations were formulated by a Guidelines Development Group (GDG) using the GRADE approach (World Health Organization, 2014). This approach entails assessing the balance of benefits and harms (at individual and population levels), certainty of evidence, values and preferences of patients and healthcare providers, resource requirements, cost-effectiveness, feasibility, and impact on equity of different interventions. Existing evidence was rigorously assessed through systematic application of GRADE methods (Schünemann et al., 2013). While all the GRADE domains were considered, there was an additional concern: falsified data. The risk of falsified data has become a critical consideration when assessing the evidence for guideline recommendations. This is because there has been a rise in the number of retracted articles in reproductive medicine (Minetto et al., 2023), as well as other biomedical fields (Hesselmann et al., 2017). Biomedical article retractions have quadrupled over the last two decades, from 10.7 to 44.8 per 100 000 publications (Freijedo-Farinas et al., 2024). Across several scientific disciplines, an estimated 1.97% of scientists have admitted fabricating, falsifying or modifying data or results at least once (Fanelli, 2009). Many reasons contribute to data fabrication and other scientific misconduct, including weak governance, incentivization of publication for career progression (Mol and Ioannidis, 2023; Freijedo-Farinas et al., 2024), so-called research paper mills (Freijedo-Farinas et al., 2024), and potential use of generative artificial intelligence, among others. Regardless of its cause, falsified data can end up being included in systematic reviews and meta-analyses (Xu et al., 2025), result in misleading or potentially harmful recommendations, and perversely influence clinical practice (Zarychanski et al., 2013; Kumar et al., 2014). Although a range of scientific misconduct and inadvertent errors can affect studies, the possibility of intentionally fraudulent data complicates the appraisal of the evidence for global guidelines, as it can influence how the GDG interprets the benefits and harms of an intervention. Inclusion or exclusion of falsified studies can shift the balance between benefits and harms of an intervention and change the strength or direction of a recommendation (Xu et al., 2025). Given these concerns, a search was conducted in the Retraction Watch Database (https://retractiondatabase.org/; The Center for Scientific Integrity, 2025) for studies included in the systematic reviews that informed the guideline recommendations. Identifying the reason for a study’s retraction is relevant in determining its potential impact on published guideline recommendations. Reasons for retractions can range from plagiarism, duplication, falsification or fabrication of data, author disputes, ethical violations (e.g. conflict of interest, missing informed consent, missing ethical approval, and compromised peer review) and other, often undocumented reasons (Stretton et al., 2012; Hesselmann et al., 2017; Gaudino et al., 2021, Minetto et al., 2023). Additionally, notices of misconduct or errors may be communicated through retractions, errata, corrections or expressions of concern (Hesselmann et al., 2017). Taking a moderate approach, if a study was retracted, had an expression of concern, or was under investigation, it was excluded from the analyses that provided the effect estimates that the GDG used to inform such recommendations (World Health Organization, 2025). Despite these actions, it is challenging to detect all falsified data unless an article has been earmarked for investigation or has been retracted, and a clear public notice issued by the journal or publisher. As others have noted, it may not be possible to capture the complete picture of retractions in the field of infertility due to the impracticality of screening the entire pool of published studies in the field (Minetto et al., 2023), and the same could be the case in other fields. The identification of fraudulent data is an ongoing process, and additional studies may come under investigation, receive an expression of concern, or be retracted after guideline publication. Therefore, WHO will continue to monitor the evidence used periodically to ensure that the evidence underlying the guideline remains valid. If a landmark article that provided an important amount of evidence to calculate estimates of effects on which recommendations were based were to be retracted, it could warrant revision of the affected recommendation. There is a lack of consensus on how to treat articles in which the lead author has another article that has already been retracted or is under investigation (Hesselmann et al., 2017; Mol and Ioannidis, 2023). While it is possible that fraudulent data could be reported in multiple articles, the guideline only excluded articles that were individually retracted or were under investigations. Other articles in which affected authors appeared were not excluded; however, a sensitivity analysis conducted showed that exclusion of these articles would not have affected the direction of effects. A notable pattern is that all the affected studies that were excluded in the guideline were randomized clinical trials (RCTs). This is reflective of the field as RCTs are the most retracted study type in the field of medically assisted reproduction (Minetto et al., 2023); however, the reliance on RCTs and reviews of RCTs in the hierarchical approach to evidence synthesis for the guideline implies a heightened risk of faulty recommendations (Minetto et al., 2023). Although examining all research protocols and individual patient data from an RCTs can provide important clues of falsified data (Mol and Ioannidis, 2023), it is not always feasible for a GDG to routinely do so for numerous recommendation questions, and we believe our approach described above mitigates the risk and reasonably safeguards the guideline. While attempts could be made to continually identify and weed fraudulent studies from guideline development processes, it is essential to address root causes of scientific misconduct not only in reproductive medicine but also more widely, through a preventive approach. In this regard, multiple stakeholders have a role in safeguarding integrity of research (Minetto et al., 2023; Mol and Ioannidis, 2023), including researchers, peer reviewers, editors, editorial boards, publishers, policy makers who can make relevant research governance regulations, governments, funding bodies, academic and research training institutions, professional societies, patient advocates, bibliographic database managers, the pharmaceutical industry, and indeed the entire scientific community. While retractions can serve as a corrective process, mitigation of scientific misconduct ought to focus on both the actors and processes that lead to it (Hesselmann et al., 2017). Although multiple mechanisms, tools and processes to reinforce research integrity exist (such as plagiarism-detecting software, retraction guidelines, research integrity or trustworthiness assessment tools, among others) they are insufficient without the individual and collective commitment from every stakeholder to uphold scientific integrity, protect patients’ interests, and advance global health. In-order to safeguard global guidelines, it is essential to institutionalize how guideline development processes – especially systematic reviewers – identify, exclude, or otherwise minimize the impact of scientific misconduct on evidence-based medicine. Transparent documentation of retraction management, including detailed reporting in the systematic review study selection (e.g. PRISMA) flow chart, could be particularly useful. Conducting searches for retracted or under-investigation studies and excluding them from the evidence that inform recommendations increases confidence in guideline recommendations. This methodological approach should be widely adopted in all WHO guideline development processes. Authors thank all individuals involved in the development of the guideline. Please see full acknowledgement statements in the guideline. The guideline was supported by the UNDP-UNFPA-UNICEF-WHO-World Bank Special Programme of Research, Development and Research Training in Human Reproduction (HRP), a cosponsored programme executed by the World Health Organization (WHO). All authors were involved in the planning or convening of the GDG. The GDG made the recommendations. The authors alone are responsible for the views expressed in this commentary, which do not necessarily represent the views, decisions, or policies of the institutions with which they are affiliated. G.M. drafted and revised the manuscript. N.S., R.B.-P., C.F., R.K., and J.K. reviewed and provided critical inputs. All authors approved the final version. Authors thank Nathan Ford for comments on an earlier draft. No specific funding was received for this commentary. All authors have no competing interests to declare.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.764
Threshold uncertainty score0.512

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.114
GPT teacher head0.508
Teacher spread0.394 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueHuman ReproductionSame topicPharmaceutical Quality and CounterfeitingFrench-language works237,207