Variation in bleeding risk estimates among online calculators
Bibliographic record
Abstract
Objective To assess the variation in bleeding risk estimates and risk stratification among Web and mobile applications for patients with atrial fibrillation. Design Cross-sectional study. Setting Simulated patient population. Participants Hypothetical patient cohorts that encompassed all possible binary risk factor combinations for each clinical prediction model. Interventions Twenty-five bleeding risk calculators (18 Web and 7 mobile apps), each of which used 1 of 4 clinical prediction models to predict an individual’s 12-month bleed risk: ATRIA (Anticoagulation and Risk Factors in Atrial Fibrillation), HAS-BLED (hypertension [systolic blood pressure >160 mm Hg], abnormal renal or liver function, stroke [caused by bleeding], bleeding, labile international normalized ratio, elderly [age >65 years], drugs [acetylsalicylic acid or nonsteroidal anti-inflammatory drugs] or alcohol [≥8 drinks per week]), HEMORR2HAGES (hepatic or renal disease, ethanol abuse, malignancy, older [age >75 years], reduced platelet count or function, rebleeding risk [history of past bleeding], hypertension [uncontrolled], anemia, genetic factors, excessive fall risk, and stroke), and mOBRI (modified Outpatient Bleeding Risk Index). Main outcome measures Four simulated cohorts were constructed. The coefficient of variation, relative difference (RD), and 95% CI for annual bleeding risk estimates were calculated for all hypothetical patient cohorts. Additionally, pairwise agreement between calculators across low- (<10%), moderate- (10% to 20%), and high-risk (>20%) categories of patients was determined. Results The risk estimates the calculators generated were imprecise, with coefficients of variation ranging from 14% for HEMORR2HAGES to 64% for mOBRI. Wide variation was observed in annual risk estimates for calculators using the mOBRI (maximum RD=4.3) and HAS-BLED (maximum RD=3.1) models. The 95% CI of mean annual bleeding risk varied among models; 1 calculator using the HAS-BLED model had a 95% CI of mean annual risk estimates of 5.4% to 6.2%, while another HAS-BLED calculator reported a 95% CI of 17.7% to 18.5%. Concordance for risk category stratification among calculators was high for those based on mOBRI and ATRIA (=1 for both). Poor agreement was observed in 1 calculator using HEMORR2HAGES (=0.54) and another using HAS-BLED ( range=-0.11 to 0.35). Conclusion Inconsistencies and a lack of precision were observed in annual risk estimates and risk stratification produced by Web and mobile bleeding risk calculators for patients with atrial fibrillation. Clinicians should refer to annual bleeding risks observed in major randomized controlled trials to inform risk estimates communicated to patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.053 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".