Predicting post-stroke cognitive impairments from lesion topography using machine learning
Bibliographic record
Abstract
Background: Stroke is the fourth and fifth leading cause of death in Canada and the United States.Survivors of stroke live with mild to severe life-long impairments.Early rehabilitation can improve long-term outcomes of stroke patients and improve their quality of life.Accurate prediction of post-stroke cognitive impairments at an individual patient level may aid the development of personalized treatments and intervention strategies. Methods:We applied and benchmarked machine learning methods on a relatively large stroke dataset (n=1401) to predict cognitive outcomes from lesion topography.The dataset included MRIs (Structural axial T1, T2-weighted spin echo, DWI and FLAIR sequence) of ischemic stroke patients carried out within 7 days from the onset of stroke and their neuropsychological assessments including measures for global cognition, language, memory, visuospatial functioning, information processing speed and executive functioning at 3 months.Three approaches to analyzing brain-behavior relationships from a predictive analytics standpoint were explored and compared in terms of out-of-sample prediction performance of post-stroke cognitive functions based on 5-fold nested crossvalidation: 1) multi-outcome models vs single-outcome models; 2) non-linear models vs linear models; and 3) data augmentation (Mixup). Results:The out-of-sample coefficient of determination (r-square) values in all approaches are generally low and inconsistent across cross-validation folds indicating poor predictive performance.However, we see that: 1) joint modeling of interrelated cognitive functions exhibits potential to perform more accurate predictions in the domains of global cognition and language; 2) non-linear models could potentially be exploited to improve individualized predictions in the domains of language and memory; and 3) it is not easy to exploit artificial samples generated by Mixup to improve the predictive performance of cognitive functions post-stroke.Résumé Contexte: L'AVC est la quatrième et la cinquième cause de décès en importance au Canada et aux États-Unis.Les survivants d'un AVC vivent avec des déficiences légères à sévères à vie.Une rééducation précoce peut améliorer les résultats à long terme des patients victimes d'un AVC et améliorer leur qualité de vie.Une prédiction précise des déficiences cognitives post-AVC au niveau d'un patient individuel peut aider au développement de traitements et de stratégies d'intervention personnalisés.Méthodes: Nous avons appliqué et comparé des méthodes d'apprentissage automatique sur un échantillon plutôt large de données d'AVC (n = 1401) pour prédire les résultats cognitifs à partir de la topographie des lésions.La banque de données comprenait des images IRM (structurelles axiales T1, spin écho pondérées en T2, DWI et séquence FLAIR) de patients victimes d'un AVC ischémique réalisées dans les 7 jours suivant le début de l'AVC et leurs évaluations neuropsychologiques, y compris des mesures de la cognition globale, du langage, de la mémoire, du fonctionnement visuospatial , dela vitesse de traitement de l'information et de la fonction exécutive à 3 mois.Trois approches pour analyser les relations cerveau-comportement du point de vue de l'analyse prédictive ont été explorées et comparées en termes de performances de prédiction hors échantillon des fonctions cognitives post-AVC, basées sur une validation croisée imbriquée 5 fois: 1) modèles multi-variés vs modèles univariés; 2) modèles non linéaires vs modèles linéaires 3) augmentation des données (Mixup).Résultats: Les valeurs du coefficient de détermination hors échantillon (rcarré) dans toutes les approches sont généralement faibles et incohérentes entre les plis de validation croisée, ce qui indique une performance prédictive médiocre.Cependant, nous voyons que: 1) la modélisation conjointe de fonctions cognitives interdépendantes présente le potentiel d'effectuer des prédictions plus précises dans les domaines de la cognition globale et du langage; 2) les modèles non Contents Contents List of Figures A.2 Mean in-sample coefficient of determination (r-square) values of various single-output linear type and non-linear models.The error bars indicate the standard deviation of the r-square values across 5 folds for each model and cognitive domain.A line at R 2 = 0.20 is drawn for the ease of visualization.56 A.3 Mean in-sample coefficient of determination (r-square) values of the Ridge and Random forest regression model.In each figure, the left panel shows results without mixup i.e., no data augmentation, the middle panel shows results with mixup with 5x data augmentation and the right panel shows results with mixup with 10x data augmentation.The error bars indicate the standard deviation of the r-square values across 5 folds for each model and cognitive domain.A line at R 2 = 0.20 is drawn for the ease of visualization.57
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".