Predicting in-hospital mortality in patients with cirrhosis: Results differ across risk adjustment methods #
Bibliographic record
Abstract
UNLABELLED: Risk-adjusted health outcomes are often used to measure the quality of hospital care, yet the optimal approach in patients with liver disease is unclear. We sought to determine whether assessments of illness severity, defined as risk for in-hospital mortality, vary across methods in patients with cirrhosis. We identified 258,731 patients with cirrhosis hospitalized in the Nationwide Inpatient Sample between 2002 and 2005. The performance of four common risk adjustment methods (the Charlson/Deyo and Elixhauser comorbidity algorithms, Disease Staging, and All Patient Refined Diagnosis Related Groups [APR-DRGs]) for predicting in-hospital mortality was determined using the c-statistic. Subgroup analyses were conducted according to a primary versus secondary diagnosis of cirrhosis and in homogeneous patient subgroups (hepatic encephalopathy, hepatocellular carcinoma, congestive heart failure, pneumonia, hip fracture, and cholelithiasis). Patients were also ranked according to the probability of death as predicted by each method, and rankings were compared across methods. Predicted mortality according to the risk adjustment methods agreed for only 55%-67% of patients. Similarly, performance of the methods for predicting in-hospital mortality varied significantly. Overall, the c-statistics (95% confidence interval) for the Charlson/Deyo and Elixhauser algorithms, Disease Staging, and APR-DRGs were 0.683 (0.680-0.687), 0.749 (0.746-0.752), 0.832 (0.829-0.834), and 0.875 (0.873-0.878), respectively. Results were robust across diagnostic subgroups, but performance was lower in patients with a primary versus secondary diagnosis of cirrhosis. CONCLUSION: Mortality analyses in patients with cirrhosis require sensitivity to the method of risk adjustment. Because different methods often produce divergent severity rankings, analyses of provider-specific outcomes may be biased depending on the method used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".