Comparison of Risk Scores for Lower Gastrointestinal Bleeding
Bibliographic record
Abstract
Importance: Clinical prediction models, or risk scores, can be used to risk stratify patients with lower gastrointestinal bleeding (LGIB), although the most discriminative score is unknown. Objective: To identify all LGIB risk scores available and compare their prognostic performance. Data Sources: A systematic search of Ovid MEDLINE, Embase, and the Cochrane Central Register of Controlled Trials from January 1, 1990, through August 31, 2021, was conducted. Non-English-language articles were excluded. Study Selection: Observational and interventional studies deriving or validating an LGIB risk score for the prediction of a clinical outcome were included. Studies including patients younger than 16 years or limited to a specific patient population or a specific cause of bleeding were excluded. Two investigators independently screened the studies, and disagreements were resolved by consensus. Data Extraction and Synthesis: Data were abstracted according to the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guideline independently by 2 investigators and pooled using random-effects models. Main Outcomes and Measures: Summary diagnostic performance measures (sensitivity, specificity, and area under the receiver operating characteristic curve [AUROC]) determined a priori were calculated for each risk score and outcome combination. Results: A total of 3268 citations were identified, of which 9 studies encompassing 12 independent cohorts and 4 risk scores (Oakland, Strate, NOBLADS [nonsteroidal anti-inflammatory drug use, no diarrhea, no abdominal tenderness, blood pressure ≤100 mm Hg, antiplatelet drug use (nonaspirin), albumin <3.0 g/dL, disease score ≥2 (according to the Charlson Comorbidity Index), and syncope], and BLEED [ongoing bleeding, low systolic blood pressure, elevated prothrombin time, erratic mental status, and unstable comorbid disease]) were included in the meta-analysis. For the prediction of safe discharge, the AUROC for the Oakland score was 0.86 (95% CI, 0.82-0.88). For major bleeding, the AUROC was 0.93 (95% CI, 0.90-0.95) for the Oakland score, 0.73 (95% CI, 0.69-0.77) for the Strate score, 0.58 (95% CI, 0.53-0.62) for the NOBLADS score, and 0.65 (95% CI, 0.61-0.69) for the BLEED score. For transfusion, the AUROC was 0.99 (95% CI, 0.98-1.00) for the Oakland score and 0.88 (95% CI, 0.85-0.90) for the NOBLADS score. For hemostasis, the AUROC was 0.36 (95% CI, 0.32-0.40) for the Oakland score, 0.82 (95% CI, 0.79-0.85) for the Strate score, and 0.24 (95% CI, 0.20-0.28) for the NOBLADS score. Conclusions and Relevance: The Oakland score was the most discriminative LGIB risk score for predicting safe discharge, major bleeding, and need for transfusion, whereas the Strate score was best for predicting need for hemostasis. This study suggests that these scores can be used to predict outcomes from LGIB and guide clinical care accordingly.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".