MétaCan
Menu
← Retour à la cohorte
Enregistrement W4404771219 · doi:10.1101/2024.11.25.24316905

Assessing the feasibility and impact of clinical trial trustworthiness checks via an application to Cochrane Reviews: Stage 2 of the INSPECT-SR project

2024· preprint· en· W4404771219 sur OpenAlexaff
Jack Wilkinson, Calvin Heal, Γεώργιος Αντωνίου, Ella Flemyng, Love Ahnström, Alessandra Alteri, Alison Avenell, Timothy Hugh Barker, David N. Borg, Nicholas J. L. Brown, Robert Buhmann, Jose Andrés Calvache, Rickard Carlsson, Lesley‐Anne Carter, Aidan G Cashin, Sarah Cotterill, Kenneth Färnqvist, Michael C Ferraro, Steph Grohmann, Lyle C. Gurrin, Jill A. Hayden, Kylie E Hunter, Natalie Hyltse, Ashma Krishan, Silvy Laporte, Toby J Lasserson, David Ruben Teindl Laursen, Sarah Lensen, Wentao Li, Tianjing Li, Jianping Liu, Clara Locher, Zuhong Lu, Andreas Lundh, Antonia Marsden, Gideon Meyerowitz‐Katz, Ben Willem Mol, Zachary Munn, Florian Naudet, David Nunan, Neil E O’Connell, Natasha Olsson, Lisa Parker, Eleftheria Patetsini, Barbara K. Redman, Sarah E.V. Rhodes, Rachel Richardson, Martin Ringsten, Ewelina Rogozińska, Anna Lene Seidler, Kyle Sheldrick, Katie Stocking, Emma Sydenham, Hugh Thomas, Sofia Tsokani, Constant Vinatier, Colby J. Vorland, Rui Wang, Bassel H. Al Wattar, Florencia Weber, Stephanie Weibel, Madelon van Wely, Chang Xu, Lisa Bero, Jamie J Kirkham

Notice bibliographique

RevuemedRxiv · 2024
Typepreprint
Langueen
DomaineDecision Sciences
ThématiqueMeta-analysis and systematic reviews
Établissements canadiensDalhousie UniversityCentre for Global Health Research
Organismes subventionnairesNational Institute for Health and Care Research
Mots-clésTrustworthinessStage (stratigraphy)Systematic reviewCochrane collaborationProcess managementComputer scienceBusinessPolitical scienceMEDLINEComputer securityLawGeology

Résumé

récupéré en direct d'OpenAlex

Abstract Background The aim of the INSPECT-SR project is to develop a tool to identify problematic RCTs in systematic reviews. In Stage 1 of the project, a list of potential trustworthiness checks was created. The checks on this list must be evaluated to determine which should be included in the INSPECT-SR tool. Methods We attempted to apply 72 trustworthiness checks to RCTs in 50 Cochrane Reviews. For each, we recorded whether the check was passed, failed or possibly failed, or whether it was not feasible to complete the check. Following application of the checks, we recorded whether we had concerns about the authenticity of each RCT. We repeated each meta-analysis after removing RCTs flagged by each check, and again after removing RCTs where we had concerns about authenticity, to estimate the impact of trustworthiness assessment. Trustworthiness assessments were compared to Risk of Bias and GRADE assessments in the reviews. Results 95 RCTs were assessed. Following application of the checks, assessors had some or serious concerns about the authenticity of 25% and 6% of the RCTs, respectively. Removing RCTs with either some or serious concerns resulted in 22% of meta-analyses having no remaining RCTs. However, many checks proved difficult to understand or implement, which may have led to unwarranted scepticism in some instances. Furthermore, we restricted assessment to meta-analyses with no more than 5 RCTs, which will distort the impact on results. No relationship was identified between trustworthiness assessment and Risk of Bias or GRADE. Conclusions This study supports the case for routine trustworthiness assessment in systematic reviews, as problematic studies do not appear to be flagged by Risk of Bias assessment. The study produced evidence on the feasibility and impact of trustworthiness checks. These results will be used, in conjunction with those from a subsequent Delphi process, to determine which checks should be included in the INSPECT-SR tool. Plain language summary Systematic reviews collate evidence from randomised controlled trials (RCTs) to find out whether health interventions are safe and effective. However, it is now recognised that the findings of some RCTs are not genuine, and some of these studies appear to have been fabricated. Various checks for these “problematic” RCTs have been proposed, but it is necessary to evaluate these checks to find out which are useful and which are feasible. We applied a comprehensive list of “trustworthiness checks” to 95 RCTs in 50 systematic reviews to learn more about them, and to see how often performing the checks would lead us to classify RCTs as being potentially inauthentic. We found that applying the checks led to concerns about the authenticity of around 1 in 3 RCTs. However, we found that many of the checks were difficult to perform and could have been misinterpreted. This might have led us to be overly sceptical in some cases. The findings from this study will be used, alongside other evidence, to decide which of these checks should be performed routinely to try to identify problematic RCTs, to stop them from being mistaken for genuine studies and potentially being used to inform healthcare decisions. What is new An extensive list of potential checks for assessing study trustworthiness was assessed via an application to 95 randomised controlled trials (RCTs) in 50 Cochrane Reviews. Following application of the checks, assessors had concerns about the authenticity of 32% of the RCTs. If these RCTs were excluded, 22% of meta-analyses would have no remaining RCTs. However, the study showed that some checks were frequently infeasible, and others could be easily misunderstood or misinterpreted. The study restricted assessment to meta-analyses including five or fewer RCTs, which might distort the impact of applying the checks.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,870
score de la tête « metaresearch » (Gemma)0,955
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesMétarecherche
DomaineSignal candidat: Évaluation · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,130
Score d'incertitude au seuil0,160

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,8700,955
Méta-épidémiologie (sens strict)0,0060,010
Méta-épidémiologie (sens large)0,0130,026
Bibliométrie0,0300,025
Études des sciences et des technologies0,0050,009
Communication savante0,0160,015
Science ouverte0,0100,021
Intégrité de la recherche0,0090,012
Charge utile insuffisante (le modèle a refusé de juger)0,0180,005

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,792
Tête enseignante GPT0,658
Écart entre enseignants0,135 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.

Devis d'étudeObservationnel
DomaineÉvaluation
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations2
Publié2024
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revuemedRxiv→Même sujetMeta-analysis and systematic reviews→Travaux en français237 207→