Abstract 23: Patient Level Prediction Models Developed With Large Observational Databases Outperformed Existing Clinical Prediction Scores For Stratifying Bleeding Risk In Patients With Atrial Fibrillation
Bibliographic record
Abstract
Background: Vitamin K antagonists (e.g., warfarin) and direct oral anticoagulants (DOACs; e.g., rivaroxaban, apixaban, dabigatran, edoxaban) are effective therapies for lowering the risk of thrombotic outcomes including stroke in patients with atrial fibrillation (AF). Like all anticoagulants, they are associated with an increased risk of bleeding. Several models (e.g., ATRIA, ORBIT, HAS-BLED, CHADS 2 and CHA 2 DS 2 -VASc) are available to predict the risk of major bleeding but their performance has not been widely evaluated in observational databases. Objective: To develop patient-level bleeding risk prediction models and evaluate their performance compared to existing clinical models of bleeding. Methods: Based on tools available through the Observational Health Data Science and Informatics Collaborative and Observed Medical Outcomes Partnership Common Data Model framework, we developed patient-level prediction (PLP) models to predict the risk of major bleeding events in patients with non-valvular AF who were new users of warfarin or DOACs. A LASSO regularized logistic regression technique including over 100,000 baseline covariates was performed across each of 4 US databases: IBM MarketScan ® (Commercial (CCAE), Medicaid (MDCD), and Medicare Supplemental (MDCR)) and Optum Clinformatics ® Extended Data Mart (Optum). We evaluated the performance of the PLP models compared with the existing models using the Area Under the Curve (AUC), with 75% training and 25% testing sets. The models were then externally validated by applying them to the other three databases. Results: In each database, all the PLP models achieved higher internal validity (AUCs 0.69 - 0.83) compared with the existing clinical models. The highest AUC among the existing clinical models was 0.76 for CHA 2 DS 2 -VASc run on the DOACs new user population in the CCAE database; the comparative AUC for the PLP model was 0.79. The external validation of the PLP models was somewhat lower, the lowest being the new user DOACs models learned on the MDCD database and applied to the other three databases with AUCs between 0.56 and 0.58. This is possibly in part due to differences between patients in MDCD versus other databases (e.g., age, disability, socioeconomic status). The highest performing was the new user warfarin model learned on the Optum database validated on the CCAE database with an AUC of 0.75. Conclusion: A patient level prediction model outperforms many existing clinical bleeding risk models currently in use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".