165 When should systematic reviews be replicated and when is it wasteful: a checklist and framework
Notice bibliographique
Résumé
In my postdoctoral research, I have used an evidence-driven, transparent process and implemented consensus approaches to develop value-added guidance on when and how to replicate systematic reviews. As outlined below, the aims, methodological approach and dissemination strategies of this project align closely with the EBM manifesto of making evidence relevant, replicable, and accessible to end-users. Background Replication is a cornerstone of the scientific method, yet replication of systematic reviews is too often overlooked, done unnecessarily or done poorly. Systematic review replication is conducted with the objective of testing whether results of an index review can be repeated or extended. Failure to replicate may lead to continued uncertainty about the implications of a body of evidence. The compelling case for replicating systematic reviews is complicated by concerns about research waste - too frequent replication of systematic reviews can represent an inefficient use of scarce research resources. There is a lack of guidance for when to, and when not to replicate systematic reviews. Objective To develop evidence-driven, consensus-based recommendations on when and how to replicate systematic reviews, taking into account the needs and preferences of the various stakeholder groups. METHODS: We used an integrated knowledge translation approach by involving an international multidisciplinary team of methodologists and knowledge users (authors, commissioners, funders, and consumers of systematic reviews, including patients, clinicians, and representatives from organizations involved with policy-making) at every stage of this research. The project was conducted in 4 phases: 1) semi-structured interviews with key informants to seek their opinions on definitions and criteria for systematic review replication; 2) a systematic review of evidence on when and how to replicate systematic reviews and an analysis of discordant reviews; 3) an online survey of knowledge users to assess level of agreement on draft criteria for systematic review replication; 4) a consensus meeting of 36 participants representing key stakeholder groups: patients, clinicians, journal editors, researchers, systematic review organizations, and guideline developers, to discuss the findings of the first three phases of the project and seek agreement on a checklist and framework for systematic review replication. Results Based on the opinion-gathering, literature review, and consensus meeting discussions, we developed: 1) a 4-item checklist applying the value of information (VOI) concept to determine whether the benefits of replicating an existing systematic review outweigh alternative uses of resources; and 2) a framework to determine how issues in the conduct of an index review represent threats to validity sufficient to justify formal replication and what are the appropriate review methods to address the specific threat of validity within the replicated systematic review. Conclusions Given the role of systematic reviews in policy-making and guideline development, the validity and reliability of their findings should be tested. The checklist and framework serve as explicit prompts to carefully consider the value of systematic review replication. Next steps will include assessing usability and acceptability of the checklist and framework, and adapting them to different users.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,760 | 0,781 |
| Méta-épidémiologie (sens strict) | 0,007 | 0,008 |
| Méta-épidémiologie (sens large) | 0,011 | 0,014 |
| Bibliométrie | 0,032 | 0,023 |
| Études des sciences et des technologies | 0,018 | 0,046 |
| Communication savante | 0,040 | 0,042 |
| Science ouverte | 0,024 | 0,028 |
| Intégrité de la recherche | 0,032 | 0,032 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,005 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».