Validation of The Umbrella Collaboration for Tertiary Evidence Synthesis in Geriatrics: Mixed Methods Study
Notice bibliographique
Résumé
BACKGROUND: The synthesis of evidence in healthcare is essential for informed decision-making and policy development. This study aims to validate The Umbrella Collaboration® (TU®), an innovative, semi-automatic tertiary evidence synthesis methodology, by comparing it with Traditional Umbrella Reviews (TUR), which are currently the gold standard. OBJECTIVE: The primary objective of this study is to evaluate whether TU®, an AI-assisted, software-driven system for tertiary evidence synthesis, can achieve comparable effectiveness to TURs, while offering a more timely, efficient, and comprehensive approach. METHODS: This comparative study evaluated TU® against TURs across eight matched projects in geriatrics. For each selected TUR, a parallel TU® project was conducted using the same research question. Outcomes of interest (OoIs), effect sizes, certainty ratings, and execution times were systematically compared. Effect sizes were assessed both quantitatively, by transforming TUR metrics to Cohen's d and correlating them with TU®'s RTU metric, and qualitatively, through categorical classifications (trivial, small, moderate, large). Certainty levels were compared by mapping GRADE ratings and TU®'s sentiment analysis scores onto a common 0-1 scale. Execution time was measured precisely in TU®, while TUR durations were estimated from literature benchmarks. Statistical analyses included chi-squared tests and Spearman correlations. RESULTS: Eight TURs in geriatrics were matched with parallel projects using TU®. TU® replicated 84.9% (73/86) of the OoIs identified by TURs and reported an additional 337 OoIs, representing a 4.77-fold increase in outcome identification. In the comparison of effect size classifications, full concordance was observed in 50.0% of cases and consistent concordance (full plus one-level deviation) in 93.8%, with a moderate strength of association (Cramér's V = 0.339). The correlation of transformed certainty values between TU® and GRADE yielded a statistically significant Spearman coefficient (ρ = 0.446; P = .025). The average execution time per TU® project was 4 hours and 46 minutes, compared to estimated durations of 6-12 months for TURs. CONCLUSIONS: The Umbrella Collaboration® demonstrated high concordance with TURs, replicating 84.9% of the outcomes identified by TURs and identifying nearly five times as many additional outcomes. The experimental effect size metric (RTU) showed moderate agreement with conventional measures, and the certainty ratings derived from sentiment analysis correlated acceptably with GRADE-based assessments. While further validation is needed, TU® appears to be a valid and efficient approach for tertiary evidence synthesis, offering a scalable and time-efficient alternative when rapid results are required. INTERNATIONAL REGISTERED REPORT: RR2-10.2196/67248.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,525 | 0,716 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,004 | 0,008 |
| Bibliométrie | 0,011 | 0,011 |
| Études des sciences et des technologies | 0,003 | 0,002 |
| Communication savante | 0,007 | 0,004 |
| Science ouverte | 0,004 | 0,008 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».