Explainable machine learning and social determinants of health in stroke prediction
Notice bibliographique
Résumé
Background: Current stroke prediction models, relying solely on traditional medical data, overlook the role of Social Determinants of Health (SDoH) like socioeconomic status and education. This narrow focus can lead to inaccurate predictions, potentially exacerbating healthcare disparities and hindering the development of effective preventive measures. This work investigates the role of SDoH in stroke and how incorporating SDoH data into AI models can improve stroke prediction, ultimately empowering healthcare providers with a more holistic view of patient risk for better decision-making and equitable healthcare delivery.Research Objectives:1. To improve the performance of stroke prediction AI models by integrating SDoH into these models.2. To ensure transparency and interpretability in stroke prediction through the application of explainable AI (XAI) methodologies.Method: The study employs datasets from the Institut de la statistique du Québec that include both clinical indicators (e.g. diabetes, heart disease, weight) and SDoH (e.g.economic, neighbourhood conditions). We applied seven machine learning models (Random Forest), Gradient Boosting Machine (GBM), CatBoost (CB), XGBoost (XGB), Light Gradient Boosting Machine (LGBM), Neural Networks (NN), and K-Nearest Neighbors (KNN) alongside XAI techniques to investigate the role SDoH plays in the models’ predictive performances. XAI methods such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) were implemented, shedding light on the influence of SDoH in the algorithms’ predictions. Performance of models was evaluated using standard metrics such as accuracy, precision, recall, F1 score and AUC (Area under the curve).Results: Our study investigated the impact of incorporating SDoH data into stroke prediction models. SDoH data variably improved performance depending on the model and specific SDoH factors incorporated, illustrating its important role alongside traditional medical data in assessing stroke risk. Our LGBM model showed maximum improvement on incorporation of SDoH features where its accuracy improved by 11.2% (from 65.9% to 77.1%). The inclusion of demographic, economic, and personal SDoH factors were the most influential. XAI methods revealed self-perceived health and stress levels as key factors for stroke prediction, emphasizing the importance of personal well-being in stroke assessment. Notably, the Light Gradient Boosting Machine (LGBM) model achieved the best performance, demonstrating an Area Under the Curve (AUC) of 81%. This translates to accuracy of 77.6%, precision of 78.6%, recall of 75.5%, and F1 score of 77.0%, showcasing LGBM’s proficiency in handling the complex relationships within SDoH data. These findings suggest the importance and potential of SDoH-integrated AI models for improved stroke prediction.Conclusion: Our findings highlight the role of SDoH data in building accurate and equitable healthcare models. Integrating SDoH factors improve stroke prediction accuracy by 1% to 3%, and foster fairer and more comprehensive patient risk assessments by considering the broader social and environmental influences on health. Furthermore, XAI techniques provide deeper insights into how SDoH and other factors contribute to predictions, promoting transparency and interpretability in these AI-driven solutions. This transparency is essential for building trust and ensuring ethically sound decision-making in healthcare
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,010 | 0,039 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,001 | 0,002 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».