Prospective and Explicit Clinical Validation of the Ottawa Heart Failure Risk Scale, With and Without Use of Quantitative <scp>NT</scp>‐pro<scp>BNP</scp>
Bibliographic record
Abstract
OBJECTIVES: We previously developed the Ottawa Heart Failure Risk Scale (OHFRS) to assist with disposition decisions for acute heart failure patients in the emergency department (ED). We sought to prospectively evaluate the accuracy, acceptability, and potential impact of OHFRS. METHODS: This prospective observational cohort study was conducted at six tertiary hospital EDs. Patients with acute heart failure were evaluated by ED physicians for the 10 OHFRS criteria and then followed for 30 days. Quantitative NT-proBNP was measured where feasible. Serious adverse event (SAE) was defined as death within 30 days, admission to monitored unit, intubation, noninvasive ventilation, myocardial infarction, or relapse resulting in hospital admission within 14 days. RESULTS: We enrolled 1,100 patients with mean (±SD) age 77.7 (±10.7) years. SAEs occurred in 170 (15.5%) cases (19.4% if admitted and 10.2% if discharged). Compared to actual practice, using an admission threshold of OHFRS score > 1 would have increased sensitivity (71.8% vs. 91.8%) but increased admissions (57.2% vs. 77.6%). For 684 cases with NT-proBNP values, using a threshold score > 1 would have significantly increased sensitivity (69.8% vs. 95.8%) while increasing admissions (60.8% vs. 88.0%). In only 11.9% of cases did physicians indicate discomfort with use of OHFRS. CONCLUSION: Prospective clinical validation found the OHFRS tool to be highly sensitive for SAEs in acute heart failure patients, albeit with an increase in admission rates. When available, NT-proBNP values further improve sensitivity. With adequate physician training, OHFRS should help improve and standardize admission practices, diminishing both unnecessary admissions for low-risk patients and unsafe discharge decisions for high-risk patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".