Applying Administrative Data‐Based Coding Algorithms for Frailty in Patients With Cirrhosis
Bibliographic record
Abstract
Frailty is a powerful prognostic tool in cirrhosis. Claims-based frailty scores estimate the presence of frailty without the need for in-person evaluation. These algorithms have not been validated in cirrhosis. Whether they measure true frailty or perform as well as frailty in outcome prediction is unknown. We evaluated 2 claims-based frailty scores-Hospital Frailty Risk Score (HFRS) and Claims-Based Frailty Index (CFI)-in 3 prospective cohorts comprising 1100 patients with cirrhosis. We assessed differences in neuromuscular/neurocognitive capabilities in those classified as frail or nonfrail based on each score. We assessed the ability of the indexes to discriminate frailty based on the Fried Frailty Index (FFI), chair stands, activities of daily living (ADL), and falls. Finally, we compared the performance of claims-based frailty measures and physical frailty measures to predict transplant-free survival using competing risk regression and patient-reported outcomes. The CFI identified neuromuscular deficits (balance, chair stands, hip strength), whereas the HFRS only identified poor chair-stand performance. The CFI had areas under the receiver operating characteristic curve (AUROCs) for identifying frailty as measured by the FFI, ADL, and falls of 0.57, 0.60, and 0.68, respectively; similarly, the AUROCs were 0.66, 0.63, and 0.67, respectively, for the HFRS. Claims-based frailty scores were associated with poor quality of life and sleep but were outperformed by the FFI and chair stands. The HFRS, per 10-point increase (but not the CFI) predicted survival of patients in the liver transplantation (subdistribution hazard ratio [SHR], 1.08; 95% confidence interval [CI], 1.03-1.12) and non-liver transplantation cohorts (SHR, 1.13; 95% CI, 1.05-1.22). Claims-based frailty scores do not adequately associate with physical frailty but are associated with important cirrhosis-related outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".