P20 Determining Clinically Meaningful Difference in Baseline EncephalApp Stroop Values to Predict HE-Related Outcomes With Multi-Center Validation
Bibliographic record
Abstract
Background: EncephalApp Stroop is a simple method to diagnose minimal hepatic encephalopathy and is linked with overt HE (OHE), & hospitalizations. Aim: (i) define Stroop OffTime+OnTime completion time test/retest variation and baseline differences in completion time that predict increased risk of OHE/hospitalizations & (ii) validate this difference in a second cohort. Methods: 3 prospective cirrhosis cohorts were enrolled. OffTime+OnTime was used as Stroop outcome (Figure 1A). Cohort 1: Stroop at baseline from 2 centers (University+VA) then followed till OHE/hospitalization or last clinical outcome. Baseline details, co-morbidities & medications were collected. Stroop values were studied using Cox proportional hazards with OHE/hospitalization as primary outcomes. Baseline Stroop values between those who developed OHE/hospitalization sooner vs rest were compared unadjusted & adjusted for clinical variables. Cohort 2: Test/retest cohort from University+VA. Underwent Stroop twice without underlying clinical change. Cohort 3: Outpatients from 10 North American sites underwent Stroop & were followed for 3 months for OHE-related hospitalizations. Baseline adjusted/unadjusted Stroop differences were compared to cohort 1. Results: 2-center cohort: 278 patients were followed for a median of 7 (3,24 IQR) months. 16% developed OHE & 12% OHE hospitalization at a median of 6 & 3 months post-testing respectively. On Cox proportional hazards for OHE, Stroop time P=0.002, MELD-Na P<0.0001 & ascites P=0.003, were significant; similar variables (Stroop P<0.001, MELD-Na P=0.009, Ascites P=0.03 & beta-blockers P=0.04) were significant for OHE hospitalization. After adjusting, we found significant baseline Stroop differences between those that developed outcomes/not (Figure 1B). Test-retest cohort: 44 patients received Stroop twice, median of 13 (4-24) months apart without significant change in OffTime+OnTime (212.4±65.1 vs 210.44±79.9 sec, P=0.75). Multi-center cohort: 357 patients of which 14 (4%) developed 3-month OHE hospitalizations, who were more likely to have prior OHE & higher MELD. Despite cohort differences, we found similar adjusted Stroop baseline differences in patients with/without OHE development (Figure 1C). Conclusion: In this prospective study with multi-center validation, >60 second OffTime+OnTime difference on Stroop portended increased risk of OHE & related hospitalizations over median 7 months, which is higher than test/retest variations. Baseline Stroop differences may add to OHE risk prediction over simple clinical models.Figure 1.: A: Study flow chart. B: Differences in Stroop at baseline for cirrhosis and demographic variables, those who did not develop overt hepatic encephalopathy compared to developed overt encephalopathy. C: Differences in seconds of Stroop in the how developed hepatic encephalopathy related events compared to those who did not.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.027 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".