MétaCan
Menu
Retour à la cohorte
Enregistrement W4382182335 · doi:10.1148/radiol.222855

A Multicenter Assessment of Interreader Reliability of LI-RADS Version 2018 for MRI and CT

2023· article· en· W4382182335 sur OpenAlexafffund
Cheng William Hong, Victoria Chernyak, Jin‐Young Choi, Sonia Lee, Chetan Potu, Timoteo Delgado, Tanya Wolfson, Anthony Gamst, Jason Birnbaum, Rony Kampalath, Chandana Lall, James T. Lee, Joseph W. Owen, Diego A. Aguirre, Mishal Mendiratta‐Lala, Matthew S. Davenport, William R. Masch, Alexandra Roudenko, Sara Lewis, Andrea S. Kierans, Elizabeth M. Hecht, Mustafa R. Bashir, Giuseppe Brancatelli, Michael Douek, Michael A. Ohliger, An Tang, Milena Cerny, Alice Fung, Eduardo A. Costa, Michael T. Corwin, John P. McGahan, Bobby Kalb, Khaled M. Elsayes, Venkateswar R. Surabhi, Katherine Blair, Robert M. Marks, Shaun Best, Ryan Ash, Karthik Ganesan, Christopher R. Kagay, Avinash Kambadakone, Jin Wang, Irene Cruite, Bijan Bijan, Mark Goodwin, Guilherme Moura Cunha, Dorathy Tamayo-Murillo, Kathryn J. Fowler, Claude B. Sirlin

Notice bibliographique

RevueRadiology · 2023
Typearticle
Langueen
DomaineMedicine
ThématiqueRadiomics and Machine Learning in Medical Imaging
Établissements canadiensUniversité de Montréal
Organismes subventionnairesNational Institute of Biomedical Imaging and BioengineeringNational Institutes of HealthFonds de Recherche du Québec - SantéFondation de l'Association des radiologistes du QuébecRadiological Society of North America
Mots-clésMedicineIntraclass correlationMalignancyRadiologyNuclear medicineMulticenter studySurgeryInternal medicine

Résumé

récupéré en direct d'OpenAlex

Background Various limitations have impacted research evaluating reader agreement for Liver Imaging Reporting and Data System (LI-RADS). Purpose To assess reader agreement of LI-RADS in an international multicenter multireader setting using scrollable images. Materials and Methods This retrospective study used deidentified clinical multiphase CT and MRI and reports with at least one untreated observation from six institutions and three countries; only qualifying examinations were submitted. Examination dates were October 2017 to August 2018 at the coordinating center. One untreated observation per examination was randomly selected using observation identifiers, and its clinically assigned features were extracted from the report. The corresponding LI-RADS version 2018 category was computed as a rescored clinical read. Each examination was randomly assigned to two of 43 research readers who independently scored the observation. Agreement for an ordinal modified four-category LI-RADS scale (LR-1, definitely benign; LR-2, probably benign; LR-3, intermediate probability of malignancy; LR-4, probably hepatocellular carcinoma [HCC]; LR-5, definitely HCC; LR-M, probably malignant but not HCC specific; and LR-TIV, tumor in vein) was computed using intraclass correlation coefficients (ICCs). Agreement was also computed for dichotomized malignancy (LR-4, LR-5, LR-M, and LR-TIV), LR-5, and LR-M. Agreement was compared between research-versus-research reads and research-versus-clinical reads. Results The study population consisted of 484 patients (mean age, 62 years ± 10 [SD]; 156 women; 93 CT examinations, 391 MRI examinations). ICCs for ordinal LI-RADS, dichotomized malignancy, LR-5, and LR-M were 0.68 (95% CI: 0.61, 0.73), 0.63 (95% CI: 0.55, 0.70), 0.58 (95% CI: 0.50, 0.66), and 0.46 (95% CI: 0.31, 0.61) respectively. Research-versus-research reader agreement was higher than research-versus-clinical agreement for modified four-category LI-RADS (ICC, 0.68 vs 0.62, respectively; P = .03) and for dichotomized malignancy (ICC, 0.63 vs 0.53, respectively; P = .005), but not for LR-5 (P = .14) or LR-M (P = .94). Conclusion There was moderate agreement for LI-RADS version 2018 overall. For some comparisons, research-versus-research reader agreement was higher than research-versus-clinical reader agreement, indicating differences between the clinical and research environments that warrant further study. © RSNA, 2023 Supplemental material is available for this article. See also the editorials by Johnson and Galgano and Smith in this issue. An earlier incorrect version appeared online. This article was corrected on June 28, 2023.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,348
Score d'incertitude au seuil0,209

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,015
Tête enseignante GPT0,352
Écart entre enseignants0,337 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations22
Publié2023
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueRadiologyMême sujetRadiomics and Machine Learning in Medical ImagingTravaux en français237 207