The data behind the image—Deep learning and its potential impact in neuro-oncological imaging
Notice bibliographique
Résumé
Deep learning techniques have rapidly evolved in the last decade and their use in neuro-oncological imaging has increasingly been proposed and evaluated recently. In a large multi-institutional endeavor, Peng et al used deep learning methods to automatically assess tumor burden in pediatric brain tumors with leptomeningeal seeding, including high-grade gliomas and medulloblastomas.1 The assessment of tumor burden is essential for response evaluation. To volumetrically assess tumor burden requires segmentation of the tumor and its potential deposits, which is a very cumbersome step when done manually. In addition, tumor segmentation is also an important prerequisite for most neuro-oncological artificial intelligence (AI)-based solutions such as the classification of tumor types or imaging-based prognostication. Every deep learning model needs to be trained with a dataset that is “labeled,” ie, a dataset in which the data are connected to some form of reference standard or ground truth. Like a human, a deep learning algorithm first needs to learn what is correct and what is incorrect before it can work “by itself” on an independent (test) dataset. Peng et al performed various manual segmentations on their pre- and postoperative datasets (including volumes of contrast-enhancing tumor and of T2/FLAIR signal abnormality) as well as RAPNO measurements for the postoperative datasets to label their data. They split their data into training and testing datasets in a 4:1 ratio and set 20% of the training data aside for validation, which is a typical ratio for this type of project. The authors then used a 3D U-Net neural network architecture with 5 levels. U-Net is a type of convolutional neural network (CNN) that is very popular in imaging-based AI and deep learning.2 CNNs are particularly good at pattern recognition and U-Net is a commonly used option for segmentation which has shown promising results in segmenting different organs and tumor types.3 U-Net consists of largely symmetric contracting and expansive parts which lead to a characteristic U-shaped architecture, hence the name. Peng et al found an excellent agreement between the tumor volumes segmented by their 3D U-Net neural networks and various manual segmentations with very high intraclass correlation coefficients of mostly above 0.9. In addition, the deep learning-based AutoRAPNO had an excellent agreement with manual RAPNO criteria, and its repeatability was higher than manual RAPNO. The authors made their code for pre- and postprocessing, training, and AutoRAPNO as well as the four pre-trained models publicly available on GitHub (https://github.com/naddan27/AutoPNeuro; accessed September 30, 2021). This is commendable and should be the standard of practice for scientific manuscript using AI and deep learning methodology. AI and deep learning techniques lend themselves to multiple different applications in neuro-oncological imaging. They can be used for tumor segmentation, as in the study by Peng et al,1 for imaging-based prediction of molecular status,4,5 for imaging-based survival prediction,6,7 or for prediction of treatment response.8 There has been an exponential increase in publications using deep learning techniques in the last years. A PubMed search with the terms “deep learning” and “brain tumor” yields 175 publications in the year 2020, but 0 publications in the year 2014. Despite this steep rise in deep learning and other methods of AI in neurooncology and beyond, very few algorithms are actually being used in clinical practice. There is a pronounced translational gap between AI-based research and clinical practice in neuroimaging. One of the biggest challenges in data science-driven research in neurooncology and in neuroradiology will be to bridge this gap and to bring more algorithms into the clinical setting. Despite the enormous potential demonstrated in myriad studies, the vast majority of algorithms are “shelved” after publication never to see the light of clinical practice. This needs to change to avoid waste in the system. Peng et al’s work addresses this gap by tackling a clinical need and proposing the use of an AI algorithm in a clinical setting. While the AI algorithm used in this paper does not contribute to the AI field per se, it demonstrates how an off-the-shelf AI algorithm can successfully be used to address an important challenge in neurooncology and neuroradiology, namely assessing tumor burden in pediatric patients with leptomeningeal seeding tumors. One of the factors contributing to the translational gap is that data scientists and clinicians commonly speak a different professional language which may lead to (sometimes fundamental) misunderstandings. We need to learn from each other and remain in constant dialogue to appreciate the potential (and the limitations) of AI for our clinical field. Training in basic methods of data science and AI, including radiomics, machine learning, and deep learning seems desirable for the neuroradiologist and neurooncologist of the future—not necessarily for them to develop these algorithms themselves, but rather to create a fundamental understanding to judge the validity of the methodology and meaning of the results. Another important factor to facilitate the translation of AI and deep learning algorithms is the accessibility in the clinical setting. These algorithms need to be integrated into the reporting workflow for neuroimaging studies with a user-friendly interface. While traditionally most AI approaches to imaging were “black boxes,” these algorithms need to become more transparent and understandable to be clinically acceptable. This requires a deeper understanding by AI scientists on how radiologists analyze images to make a diagnosis, and how to clinically translate the AI results in a way that is understandable. Moreover, these continually learning algorithms need to be continuously quality controlled to prevent them from “going bad.” We eventually envision a “human-centered” AI in neuroimaging in which there is a human-machine partnership with the human always being in the center. This will not only provide the human with extra information necessary for an informed decision making, but it will also enable them to provide feedback to the AI system to constantly improve performance. This human-centered AI should be patient-centric by being personalized and precise and physician-centric by being accessible and transparent, while at the same time undergoing rigorous continuous quality control. The text is the sole product of the authors and no third party had input or gave support to its writing. Conflict of interest statement. None declared.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,035 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,004 |
| Communication savante | 0,003 | 0,005 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,020 | 0,029 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».