MétaCan
Menu
Back to cohort

Assessing the feasibility and impact of clinical trial trustworthiness checks via an application to Cochrane Reviews: Stage 2 of the INSPECT-SR project

2025· article· en· W4410221436 on OpenAlexaff
Jack Wilkinson, Calvin Heal, George Α. Antoniou, Ella Flemyng, Love Ahnström, Alessandra Alteri, Alison Avenell, Timothy Hugh Barker, David N. Borg, Nicholas J. L. Brown, Robert Buhmann, Jose Andrés Calvache, Rickard Carlsson, Lesley‐Anne Carter, Aidan G Cashin, Sarah Cotterill, Kenneth Färnqvist, Michael C Ferraro, Steph Grohmann, Lyle C. Gurrin, Jill A. Hayden, Kylie E Hunter, Natalie Hyltse, Ashma Krishan, Silvy Laporte, Toby J Lasserson, David Ruben Teindl Laursen, Sarah Lensen, Wentao Li, Tianjing Li, Jianping Liu, Clara Locher, Zewen Lu, Andreas Lundh, Antonia Marsden, Gideon Meyerowitz‐Katz, Ben W. Mol, Zachary Munn, Florian Naudet, David Nunan, Neil E O’Connell, Natasha Olsson, Lisa Parker, Eleftheria Patetsini, Barbara K. Redman, Sarah Rhodes, Rachel Richardson, Martin Ringsten, Ewelina Rogozińska, Anna Lene Seidler, Kyle Sheldrick, Katie Stocking, Emma Sydenham, Hugh Thomas, Sofia Tsokani, Constant Vinatier, Colby J. Vorland, Rui Wang, Bassel H. Al Wattar, Florencia Weber, Stephanie Weibel, Madelon van Wely, Chang Xu, Lisa Bero, Jamie J Kirkham

Bibliographic record

VenueJournal of Clinical Epidemiology · 2025
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsDalhousie University
FundersDepartment of Health and Social CareNational Institute for Health and Care Research
KeywordsTrustworthinessStage (stratigraphy)Systematic reviewMeta-analysisMedicineCochrane collaborationClinical trialMedical physicsMEDLINEComputer scienceInternal medicineCochrane LibraryPolitical science

Abstract

fetched live from OpenAlex

BACKGROUND AND OBJECTIVES: The aim of the INveStigating ProblEmatic Clinical Trials in Systematic Reviews (INSPECT-SR) project is to develop a tool to identify problematic RCTs in systematic reviews. In stage 1 of the project, a list of potential trustworthiness checks was created. The checks on this list must be evaluated to determine which should be included in the INSPECT-SR tool. METHODS: We attempted to apply 72 trustworthiness checks to randomized controlled trials (RCTs) in 50 Cochrane reviews. For each, we recorded whether the check was passed, failed, or possibly failed or whether it was not feasible to complete the check. Following application of the checks, we recorded whether we had concerns about the authenticity of each RCT. We repeated each meta-analysis after removing RCTs flagged by each check and again after removing RCTs where we had concerns about authenticity to estimate the impact of trustworthiness assessment. Trustworthiness assessments were compared to Risk of Bias and Grading of Recommendations Assessment, Development and Evaluation (GRADE) assessments in the reviews. RESULTS: Ninety-five RCTs were assessed. Following application of the checks, assessors had some or serious concerns about the authenticity of 25% and 6% of the RCTs, respectively. Removing RCTs with either some or serious concerns resulted in 22% of meta-analyses having no remaining RCTs. However, many checks proved difficult to understand or implement, which may have led to unwarranted skepticism in some instances. Furthermore, we restricted assessment to meta-analyses with no more than five RCTs (54% contained only 1 RCT), which will distort the impact on results. No relationship was identified between trustworthiness assessment and Risk of Bias or GRADE. CONCLUSION: This study supports the case for routine trustworthiness assessment in systematic reviews, as problematic studies do not appear to be flagged by Risk of Bias assessment. The study produced evidence on the feasibility and impact of trustworthiness checks. These results will be used, in conjunction with those from a subsequent Delphi process, to determine which checks should be included in the INSPECT-SR tool. PLAIN LANGUAGE SUMMARY: Systematic reviews collate evidence from randomized controlled trials (RCTs) to find out whether health interventions are safe and effective. However, it is now recognized that the findings of some RCTs are not genuine, and some of these studies appear to have been fabricated. Various checks for these "problematic" RCTs have been proposed, but it is necessary to evaluate these checks to find out which are useful and which are feasible. We applied a comprehensive list of "trustworthiness checks" to 95 RCTs in 50 systematic reviews to learn more about them and to see how often performing the checks would lead us to classify RCTs as being potentially inauthentic. We found that applying the checks led to concerns about the authenticity of around 1 in three RCTs. However, we found that many of the checks were difficult to perform and could have been misinterpreted. This might have led us to be overly skeptical in some cases. The findings from this study will be used, alongside other evidence, to decide which of these checks should be performed routinely to try to identify problematic RCTs, to stop them from being mistaken for genuine studies and potentially being used to inform health care decisions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.874
metaresearch head score (Gemma)0.953
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.126
Threshold uncertainty score0.155

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.8740.953
Meta-epidemiology (narrow)0.0070.011
Meta-epidemiology (broad)0.0140.028
Bibliometrics0.0350.031
Science and technology studies0.0060.010
Scholarly communication0.0170.016
Open science0.0110.023
Research integrity0.0100.011
Insufficient payload (model declined to judge)0.0160.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.916
GPT teacher head0.756
Teacher spread0.160 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations19
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Clinical EpidemiologySame topicMeta-analysis and systematic reviewsFrench-language works237,207