Development of machine learning models predicting mortality using routinely collected observational health data from 0-59 months old children admitted to an intensive care unit in Bangladesh: critical role of biochemistry and haematology data
Notice bibliographique
Résumé
Introduction Treatment in the intensive care unit (ICU) generates complex data where machine learning (ML) modelling could be beneficial. Using routine hospital data, we evaluated the ability of multiple ML models to predict inpatient mortality in a paediatric population in a low/middle-income country. Method We retrospectively analysed hospital record data from 0-59 months old children admitted to the ICU of Dhaka hospital of International Centre for Diarrhoeal Disease Research, Bangladesh. Five commonly used ML models- logistic regression, least absolute shrinkage and selection operator, elastic net, gradient boosting trees (GBT) and random forest (RF), were evaluated using the area under the receiver operating characteristic curve (AUROC). Top predictors were selected using RF mean decrease Gini scores as the feature importance values. Results Data from 5669 children was used and was reduced to 3505 patients (10% death, 90% survived) following missing data removal. The mean patient age was 10.8 months (SD=10.5). The top performing models based on the validation performance measured by mean 10-fold cross-validation AUROC on the training data set were RF and GBT. Hyperparameters were selected using cross-validation and then tested in an unseen test set. The models developed used demographic, anthropometric, clinical, biochemistry and haematological data for mortality prediction. We found RF consistently outperformed GBT and predicted the mortality with AUROC of ≥0.87 in the test set when three or more laboratory measurements were included. However, after the inclusion of a fourth laboratory measurement, very minor predictive gains (AUROC 0.87 vs 0.88) resulted. The best predictors were the biochemistry and haematological measurements, with the top predictors being total CO 2 , potassium, creatinine and total calcium. Conclusions Mortality in children admitted to ICU can be predicted with high accuracy using RF ML models in a real-life data set using multiple laboratory measurements with the most important features primarily coming from patient biochemistry and haematology.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».