Concordance with CONSORT-AI guidelines in reporting of randomised controlled trials investigating artificial intelligence in oncology: a systematic review
Notice bibliographique
Résumé
Background: The advent of artificial intelligence (AI) tools in oncology to support clinical decision-making, reduce physician workload and automate workflow inefficiencies yields both great promise and caution. To generate high-quality evidence on the safety and efficacy of AI interventions, randomised controlled trials (RCTs) remain the gold standard. However, the completeness and quality of reporting among AI trials in oncology remains unknown. Objective: This systematic review investigates the reporting concordance of RCTs for AI interventions in oncology using the CONSORT (Consolidated Standards of Reporting Trials) 2010 and CONSORT-AI 2020 extension guideline and comprehensively summarises the state of AI RCTs in oncology. Methods and analysis: We queried OVID MEDLINE and Embase on 22 October 2024 using AI, cancer and RCT search terms. Studies were included if they reported on an AI intervention in an RCT including participants with cancer. Results: This study included 57 RCTs of AI interventions in oncology that were primarily focused on screening (54%) or diagnosis (19%) and intended for clinician use (88%). Among all 57 RCTs, median concordance with CONSORT 2010 and CONSORT-AI 2020 was 82%. Compared with trials published before the release of CONSORT-AI (n=8), trials published after the release of CONSORT-AI (n=49) had lower median overall CONSORT (82% vs 92%) and CONSORT 2010 (81% vs 92%) concordance but similar CONSORT-AI median concordance (93% vs 93%). Guideline items related to study methodology necessary for reproducibility using the AI intervention, such as input data inclusion and exclusion, algorithm version, low quality data handling, assessment of performance error and data accessibility, were consistently under-reported. When stratifying included trials by their overall risk of bias, trials at serious risk of bias (57%) were less concordant to CONSORT guidelines compared with trials at moderate (71%) or low (84%) risk of bias. Conclusion: Although the majority of CONSORT and CONSORT-AI items were well-reported, critical gaps related to reporting of methodology, reproducibility and harms persist. Addressing these gaps through consideration of trial design to mitigate risks of bias coupled with standardised reporting is one step towards responsible adoption of AI to improve patient outcomes in oncology.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,054 | 0,443 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,030 | 0,001 |
| Bibliométrie | 0,001 | 0,002 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».