Comparison of bleeding risk scores in patients with atrial fibrillation: insights from the RE‐LY trial
Bibliographic record
Abstract
Abstract Background Oral anticoagulation is the mainstay of stroke prevention in atrial fibrillation (AF), but must be balanced against the associated bleeding risk. Several risk scores have been proposed for prediction of bleeding events in patients with AF. Objectives To compare the performance of contemporary clinical bleeding risk scores in 18 113 patients with AF randomized to dabigatran 110 mg, 150 mg or warfarin in the RE‐LY trial. Methods HAS‐BLED, ORBIT, ATRIA and HEMORR 2 HAGES bleeding risk scores were calculated based on clinical information at baseline. All major bleeding events were centrally adjudicated. Results There were 1182 (6.5%) major bleeding events during a median follow‐up of 2.0 years. For all the four schemes, high‐risk subgroups had higher risk of major bleeding (all P < 0.001). The ORBIT score showed the best discrimination with c‐indices of 0.66, 0.66 and 0.62, respectively, for major, life‐threatening and intracranial bleeding, which were significantly better than for the HAS‐BLED score (difference in c‐indices: 0.050, 0.053 and 0.048, respectively, all P < 0.05). The ORBIT score also showed the best calibration compared with previous data. Significant treatment interactions between the bleeding scores and the risk of major bleeding with dabigatran 150 mg BD versus warfarin were found for the ORBIT ( P = 0.0019), ATRIA ( P < 0.001) and HEMORR 2 HAGES ( P < 0.001) scores. HAS‐BLED score showed a nonsignificant trend for interaction ( P = 0.0607). Conclusions Amongst the current clinical bleeding risk scores, the ORBIT score demonstrated the best discrimination and calibration. All the scores demonstrated, to a variable extent, an interaction with bleeding risk associated with dabigatran or warfarin.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".