MétaCan
Menu
Retour à la cohorte
Enregistrement W4417515038 · doi:10.1038/s41598-025-31562-5

Predicting and classifying type 2 diabetes using a transparent ensemble model combining random forest, k-nearest neighbor, and neural networks

2025· article· en· W4417515038 sur OpenAlexaff
Niloufar Zaferani, Mohammad Reza Afrash, Khadijeh Moulaei

Notice bibliographique

RevueScientific Reports · 2025
Typearticle
Langueen
DomaineHealth Professions
ThématiqueArtificial Intelligence in Healthcare
Établissements canadiensArtificial Intelligence in Medicine (Canada)
Organismes subventionnairesnon disponible
Mots-clésInterpretabilityRandom forestEnsemble learningArtificial neural networkFeature selectionMissing dataDecision treeLeverage (statistics)Deep learningData pre-processing

Résumé

récupéré en direct d'OpenAlex

Diabetes is one of the major health challenges in today's world, since chronic elevation of blood sugar can cause serious and sometimes irreparable damage to organs such as the heart, kidneys, and nervous system. Early detection of this disease plays a vital role in reducing its complications. However, machine learning and deep learning models often face distrust in medical settings due to their opaque, "black-box" nature. The aim of this study was to combine three machine learning algorithms using stacking and voting methods to propose a model for type 2 diabetes detection, and to increase transparency by using the explainability techniques LIME and SHAP to identify important features. This study used medical data from 768 Pima Indians Diabetes samples, including 8 features such as age, BMI, glucose, insulin, blood pressure, skin thickness, pregnancies, and family history. Data preprocessing included mean imputation for missing or zero values, Min-Max normalization, and classification into "Normal", "Prediabetes", and "Diabetes" based on fasting glucose thresholds. Feature selection was performed using Spearman correlation to retain the most relevant variables. A hybrid machine learning model was developed using three base models Neural Network (NN), k-Nearest Neighbors (KNN), and Random Forest (RF) with automated hyperparameter tuning. The outputs of these models were combined via stacking using a logistic regression (LR) meta-model and in parallel using a soft voting method. Nested cross-validation (5 outer and 5 inner folds) was applied to prevent data leakage and ensure robust evaluation. Model interpretability was assessed using LIME for local explanations and SHAP for global feature importance. Decision thresholds and influential feature regions were identified, and model calibration and decision curves evaluated clinical reliability. Models performance was evaluated using accuracy, precision, recall, specificity, F1-score, AUROC, Brier Score (1-B), and Expected Calibration Error (1-E). Statistical reliability was assessed using bootstrap resampling to compute 95% confidence intervals, as well as paired tests to compare the hybrid model with the base models and voting ensemble. Based on the evaluation metrics, the stacking ensemble achieved perfect performance for Class 0, with 100% accuracy, precision, recall, specificity, F1 score, and AUROC, alongside the highest calibration metrics (Brier Score: 99.9, ECE: 98.7). The Random Forest model also excelled, achieving 100% accuracy, precision, recall, specificity, and F1 score for Class 0 and Class 2. In contrast, the KNN model consistently underperformed, particularly for Class 0 (F1: 83.3, Precision: 83.3, Recall: 83.3). The Neural Network demonstrated strong recall for Class 0 (100%), while the voting ensemble showed balanced results but was slightly outperformed by the top ensemble methods. Explainable AI analyses using LIME and SHAP revealed that glucose was the most influential predictor for identifying the Pre-diabetes state. Both methods consistently identified a decision band between 0.35 and 0.47 (corresponding to 100-125 mg/dL) as the transition zone between "Normal" and "Prediabetes", confirming the model's alignment with WHO/ADA diagnostic criteria. The stacking model achieved perfect performance and superior calibration, outperforming all other models in type 2 diabetes prediction and classification. Explainability techniques (LIME and SHAP) identified glucose level, body mass index, and blood pressure as key predictive factors. This approach provides an accurate and interpretable tool for clinical decision support in healthcare systems.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesÉtudes des sciences et des technologies
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,928
Score d'incertitude au seuil0,998

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0020,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,001
Études des sciences et des technologies0,0030,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,149
Tête enseignante GPT0,422
Écart entre enseignants0,273 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueScientific ReportsMême sujetArtificial Intelligence in HealthcareTravaux en français237 207