Optimizing personalized treatment strategies for coronary artery disease using deep learning and reinforcement learning
Bibliographic record
Abstract
Coronary artery disease (CAD) is a leading global cause of mortality and morbidity, affecting approximately 295 million people and resulting in 9.5 million deaths in 2021. Despite advancements in medical interventions that have reduced CAD mortality in high-income countries, significant challenges remain in delivering optimized, personalized care, particularly in settings characterized by high patient variability. Traditional decision-making in the treatment of obstructive CAD—encompassing percutaneous coronary intervention (PCI), coronary artery bypass grafting (CABG), and medical therapy—relies predominantly on population-level data from randomized controlled trials. However, this approach often neglects individual patient characteristics and the sequential nature of CAD treatments. This thesis employs machine learning (ML), with a focus on reinforcement learning (RL) and deep learning techniques, to improve CAD treatment decision-making. Specifically, it develops an offline RL framework designed to provide personalized treatment recommendations for patients with CAD, utilizing a large cohort from Alberta, Canada. Chapter 3 addresses the challenge of processing high-dimensional clinical data for ML applications. Through extensive experimentation with several unsupervised feature selection techniques, this study identifies the weight-adjusted concrete autoencoder as the most effective method for extracting efficient and interpretable features from diagnostic (ICD-10) and therapeutic (ATC) code databases. The selected features serve as the foundation for the analyses in subsequent chapters. Chapter 4 forms the core of this thesis, presenting an offline RL framework named RL4CAD to optimize CAD treatment recommendations tailored to individual patient profiles. Off-policy evaluation of RL4CAD models demonstrates their superiority over traditional physician-driven decision-making, achieving significant reductions in major adverse cardiovascular events. By using conservative RL models, we balanced the optimal recommendations with the current clinical practice. Moreover, this framework introduces interpretability to RL-based decisions by employing models with limited state spaces and identifying key features that influence outcomes. In Chapter 5, the thesis addresses the challenge of distribution shifts across diverse patient populations, such as variations in sex and treatment site, which complicate the optimization of CAD treatment using RL. By stratifying patient cohorts by these features and independently evaluating physician behavior and RL-derived policies, significant disparities in clinical practices were identified. To mitigate these challenges, transfer learning was integrated into the RL framework, enabling the model to adapt to diverse patient subgroups with minimal data and retraining. This approach effectively addressed distribution shifts, improving the RL models’ capacity to deliver personalized and equitable CAD treatment recommendations in patient populations that they had not been trained on. Collectively, the innovations in feature selection, RL-guided decision-making, and addressing distribution shifts contribute a scalable, data-driven solution for CAD management. By delivering personalized treatment recommendations that account for patient and practice variability, this work advances the potential of RL-driven precision medicine in cardiovascular care. Furthermore, the findings from this thesis establish a foundation for applying RL to other complex and dynamic treatment domains.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".