#3815 A DEEP LEARNING APPROACH TO PERSONALISED ANTI-HYPERTENSIVE MEDICATION TITRATION
Bibliographic record
Abstract
Abstract Background and Aims Hypertension is the number one risk factor for premature death worldwide. Artificial Intelligence (AI) Clinical Decision Support Systems are an important next step for hypertension management but require rigorous evaluation before assimilation into routine clinical practice. Our aim is to develop and evaluate an AI clinical decision support tool for hypertension management trained on randomised clinical trial data. Method The Systolic Blood Pressure Intervention Trial (SPRINT) trial showed that intensive BP control to SBP <120 mm Hg results in significant cardiovascular benefit in high-risk patients with hypertension compared with routine BP control to <140 mm Hg. We trained a feed-forward neural network using Keras and Tensorflow for R and data from 9361 persons in the SPRINT randomized clinical trial to predict the probability of increase, reduce or no-change in total number of anti-hypertensive medications at each visit. The network is designed to model SPRINT investigator deviations from the protocol. Six baseline patient variables (age, sex, race, aspirin use, eGFR, group assignment (intensive or standard)), two visit patient variables (Systolic and Diastolic Blood Pressure), and two previous visit variables (Systolic and Diastolic Blood Pressure) were inputted into the model after data normalisation and centering. Hyperparameters were tuned using a grid search method. We conducted internal validation using a 20% validation set. Clinical validation was performed using an unseen test set of typical hypertension scenarios (n = 50). An R Shiny app was developed to enter new patient information and display a sensitivity analysis of probabilities calculated after changing each baseline variable by plus and minus 10 (Figure 1). To avoid poor performance on new out of distribution cases we tested five methods for Out of Distribution (OOD) detector, 1. Autoencoders, 2. Normalizing flow, 3. Local Outlier Factor (LOF), 4. Probabilistic Principal Component Analysis (PPCA) and Kernel Density Estimation (KDE). Results The accuracy was 76.7% on the validation dataset after 1000 Epochs. Clinical validation revealed suboptimal performance for unseen Out of Distribution data (e.g. recommending reduction of medication when SBP >160 and number of medications was zero). This poor performance on OOD samples prompted the creation of an OOD detector to use in series with the AI Decision Support System for Hypertension. The AUROCs were 0.77 for AE, 0.96 for Normalizing Flow, 0.82 for LOF, 0.88 for PPCA and 0.96 for the Kernel Density Estimation. Conclusion Artificial Intelligence (AI) Clinical Decision Support Systems for Hypertension are feasible and an important next step for hypertension management but generalisability and safety require rigorous evaluation before assimilation into routine clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".