The Role of Artificial Intelligence in Predicting Optic Neuritis Subtypes From Ocular Fundus Photographs
Notice bibliographique
Résumé
BACKGROUND: Optic neuritis (ON) is a complex clinical syndrome that has diverse etiologies and treatments based on its subtypes. Notably, ON associated with multiple sclerosis (MS ON) has a good prognosis for recovery irrespective of treatment, whereas ON associated with other conditions including neuromyelitis optica spectrum disorders or myelin oligodendrocyte glycoprotein antibody-associated disease is often associated with less favorable outcomes. Delay in treatment of these non-MS ON subtypes can lead to irreversible vision loss. It is important to distinguish MS ON from other ON subtypes early, to guide appropriate management. Yet, identifying ON and differentiating subtypes can be challenging as MRI and serological antibody test results are not always readily available in the acute setting. The purpose of this study is to develop a deep learning artificial intelligence (AI) algorithm to predict subtype based on fundus photographs, to aid the diagnostic evaluation of patients with suspected ON. METHODS: This was a retrospective study of patients with ON seen at our institution between 2007 and 2022. Fundus photographs (1,599) were retrospectively collected from a total of 321 patients classified into 2 groups: MS ON (262 patients; 1,114 photographs) and non-MS ON (59 patients; 485 photographs). The dataset was divided into training and holdout test sets with an 80%/20% ratio, using stratified sampling to ensure equal representation of MS ON and non-MS ON patients in both sets. Model hyperparameters were tuned using 5-fold cross-validation on the training dataset. The overall performance and generalizability of the model was subsequently evaluated on the holdout test set. RESULTS: The receiver operating characteristic (ROC) curve for the developed model, evaluated on the holdout test dataset, yielded an area under the ROC curve of 0.83 (95% confidence interval [CI], 0.72-0.92). The model attained an accuracy of 76.2% (95% CI, 68.4-83.1), a sensitivity of 74.2% (95% CI, 55.9-87.4) and a specificity of 76.9% (95% CI, 67.6-85.0) in classifying images as non-MS-related ON. CONCLUSION: This study provides preliminary evidence supporting a role for AI in differentiating non-MS ON subtypes from MS ON. Future work will aim to increase the size of the dataset and explore the role of combining clinical and paraclinical measures to refine deep learning models over time.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,009 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».