Predicting Major Bleeding in Patients with Venous Thromboembolism on Extended Anticoagulation Therapy Using Follow-up Data and Long Short-Term Memory Based Recurrent Neural Network
Notice bibliographique
Résumé
Background: It is often challenging for physicians to decide on the duration of anticoagulation treatment for patients with venous thromboembolism (VTE), as they need to weigh the risk of recurrent thrombosis and bleeding at the same time. Several clinical models have been developed to help physicians identify patients with high risk of bleeding. However, these tools use only the baseline clinical information and are not able to incorporate the clinical conditions and events that occur over time that may influence the risk of bleeding. Therefore, capturing the patterns and relationships in the continuously changing time series of clinical data could be a more powerful approach to developing predictor models than just relying on the baseline clinical information. Nonetheless, creating predictive models from follow-up clinical information is challenging given that time series clinical data are non-uniform, high-dimensional, multivariate observations. Here, we present the first attempt at creating a model that uses the time series follow-up information to predict major bleeding over time using a Recurrent Neural Network (RNN) with Long Short-Term Memory (LSTM) cells. Methods: 2542 patients diagnosed with VTE were enrolled in a prospective cohort study over 8 years. In addition to recording their clinical information at baseline, 6-month follow-up interviews were conducted using a standard script to monitor bleeding status and record clinical information. Major bleeding was defined by the International Society on Thrombosis and Haemostasis, with suspected bleeding events classified by an independent adjudication committee. Overall, 118 patients had major bleeding - a 4.6% incidence rate. The median and mode of the clinical variables were used to impute missing numerical and categorical values, respectively, and for patients who had no follow-up information, for whom bleeding occurred before the first follow-up, an artificial follow-up data point was generated from their corresponding baseline data. Thereafter, the data was divided into two stratified sets: 70% for training, and 30% for testing. Five supervised neural network-based machine learning models with different architectures were trained on the baseline dataset, or the follow-up dataset, or both to predict major bleeding. After training, these machine learning models were tested on the testing set and compared to the conventional clinical models, modified to make them compatible with the available predictor variables in our dataset, including the CHAP, the HAS-BLED, the VTE-BLEED, the RIETE, the ACCP, and the OBRI, which only use the baseline information. Results: Overall, the models that used the follow-up information had a higher area under the Receiver Operating Curve (AUROC) or c-statistic compared to the other models that only relied on the baseline dataset. In particular, the LSTM RNN model was able to achieve AUROC of 81.3% that is more than 10% higher compared to the best performing clinical model. We discovered that the LSTM RNN model mostly relied on features such as number of concomitant medications, years since baseline visit, use of specific antibiotics or antiplatelet agents, and presence of new hypertension to predict bleeding from the follow-up dataset. Furthermore, half of the bleeding events occurred within the first year after patients' baseline visits - a trend reflected in the predictions made by LSTM RNN model. Finally, the models that used both the baseline and the follow-up datasets showed different results depending on their architectures; that is, the simpler ensemble model achieved AUROC of 82.5% while the more complex model had AUROC of 70.8% due to overfitting. Conclusion: We have shown that using time series follow-up data can improve bleeding risk prediction in patients with VTE who are on extended anticoagulant therapy compared to just using the baseline data, and clinicians might benefit from using such an approach. Furthermore, our results indicate that LSTM RNN is a suitable architecture to model routine clinical follow-up data. Finally, we believe using time series data could improve the performance of the other clinical models that are currently based on one-time baseline measurements.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».