From Deep Learning to Deep Consciousness:Predicting Comatose Patient Outcomes with Time-Series CT Scans
Notice bibliographique
Résumé
The aim of this thesis was to predict the outcome of patients based on time series CT scans using Deep Learning techniques. Previous studies that have attempted to predict patient outcomes using only CT or MRI scans have largely been unsuccessful. In contrast, this study presented a novel approach by incorporating the temporal dimension of CT scans, thereby potentially capturing dynamic changes that may be critical to predicting outcomes. Over the first two years of this study, acquiring and hosting the dataset from Copenhagen University Hospital in Denmark presented considerable hurdles. Complying with GDPR rules, ethical guidelines, and data management clearances alone required nearly eight months. After obtaining approval, we initially intended to store the data on the DTU server, which adhered to GDPR standards. However, we were only informed after several months that GPU support for sensitive data had been discontinued on that server. This prompted us to revise our data management strategy and migrate the dataset to the High-Performance Computing (HPC) server of the Clinically Applied Artificial Intelligence (CAAI) Center. This process ultimately occupied the PhD’s first two years. In the interim, multiple methodological approaches were explored and developed in anticipation of receiving the dataset, ensuring that we were fully prepared once it became available. To address the initial unavailability of data, we divided the PhD timeline into two phases: Pre-Deep Consciousness and Deep Consciousness. The Pre-Deep Consciousness phase included a series of scientific experiments aimed at developing a robust deep learning segmentation approach. Our strategy involved extracting latent space representations from the images and training time series models (LSTM model and transformer model) on these features to predict outcomes for coma patients. The first step of our Pre-Deep Consciousness plan was to validate our hypothesis by developing a synthetic dataset that mimics the time series nature of CT scans. This synthetic dataset included both images and corresponding masks. We trained a U-Net model on these synthetic images and masks, which allowed us to explore two different approaches. In the first approach, we extracted the latent space from the encoder of the U-Net and fed it into a Long Short-Term Memory (LSTM) network to predict patient outcomes. In the second approach, we extracted radiomics features from the images and used them as input for an LSTM model. Using the synthetic dataset, we achieved promising results with an AUC of approximately 0.92; however, this was based on thousands of generated images. To evaluate the performance of our methods on a smaller, real-world dataset, we also applied these approaches to the Brats-18 dataset. Unlike our synthetic dataset, Brats-18 contains only a single 3D MRI image per patient, rather than a series of scans. On this dataset, survival prediction yielded an AUC of about 0.55, highlighting the challenges of applying our methods to limited data. After concluding experiments on the synthetic and BraTS-18 datasets, our data acquisition process was still ongoing. We began investigating robust segmentation methods using a small dataset. Because we anticipated that, due to time constraints, we would only be able to create masks for fewer than 50 head CT images, we planned to use latent representations for further analysis. Building on their promising performance, we experimented with various versions of Multi-Planar UNet, the Segment Anything Model (SAM), and You Only Look Once (YOLO). For simplicity, we began with 2D images, focusing on segmenting the short-axis abdominal aorta in POCUS (Point-of-Care Ultrasound) images. With only 500 training samples, YOLOv8 segmented nearly every image, including noisy POCUS scans sourced online. We further experimented by training YOLO on just 100 samples (images and masks) and passing the resulting bounding boxes to SAM as prompts, leading to masks noticeably superior to those produced by YOLO alone—suggesting that SAM’s encoder might provide a useful latent space for our project. After concluding this phase, we searched for a place to host the Deep Consciousness dataset and explored another segmentation method: Multi-Planar UNet (MPU). This model excels at handling small, complex datasets (fewer than 100 samples), such as knee MRI scans, and when we added an attention mechanism, MPU’s Dice Score improved by about 2%. Meanwhile, during our YOLO experiments, we uncovered two major obstacles to its application in medical imaging: a complex training data structure and a lack of medical-specific pre-trained weights. In response, we open-sourced three critical resources—a Python package that simplifies training 2D and 3D images with YOLO, “Med-YOLO” pre-trained weights, and a COCO-style medical dataset containing roughly 200,000 2D CT scans and around 4.5 million bounding boxes—to help make medical imaging research more accessible and efficient. In the midst of the MED-YOLO project, we acquired our dataset and launched the Deep Consciousness project, which was divided into two stages. The preliminary assessment stage involved only 33% of patients for whom survival labels were available, focusing on those with more than three CT scans to capture temporal effects. In the final assessment stage, after we obtained all labels for all patients, we included individuals with even a single CT scan, building on the preliminary findings. The first step of this project was bias analysis and preprocessing. We checked for correlations between the number of CT scans per patient and, finding none, proceeded to generate brain masks using TotalSegmentator. Next, we calculated mask volumes to exclude cases lacking head CT information, ensuring that remaining scans were within 1,000 to 2,000 cc in size. Once these scans were identified, they were registered to the Montreal Neurological Institute (MNI) head CT, and SAM MED-3D was employed to extract latent space representations. Two deep learning models: Transformer based model and a Conv-LSTM model were trained on the time-series latent representations, achieving AUROCs of 0.7 and 0.65, respectively, in 5-fold cross-validation on the validation dataset.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,008 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».