S922 Multicenter Validation of Screening Tool for Diagnosing Non-Alcoholic Fatty Liver Disease in Patients with Crohn’s Disease
Bibliographic record
Abstract
Introduction: Crohn’s Disease (CD) patients are twice as likely compared to controls to develop nonalcoholic fatty liver disease (NAFLD) leading to increased risk of cardiometabolic complications. Given this, we developed the first clinical screening tool for NAFLD in CD, Clinical Predictor for NAFLD in CD (CPN-CD) which uses readily accessible laboratory and clinical parameters. We have demonstrated CPN-CD outperforms the Hepatic Steatosis Index in detecting NAFLD in CD in an internal cohort. Here we performed a multicenter analysis to externally validate CPN-CD against transient elastography (TE) and establish its diagnostic accuracy to detect NAFLD in patients with CD at two different controlled attenuation parameter (CAP) thresholds. Methods: A total of 454 patients with CD across four prospective cohorts from tertiary IBD centers across United States, Canada and India were included in the study. Of these, 412 patients were screened with TE only to determine prevalence of hepatic steatosis while 42 were screened with additional gold standard magnetic resonance imaging derived proton density fat fraction (MRI-PDFF), where hepatic steatosis was defined as ≥ 5.5% fat density, then reflexed to TE. To evaluate discriminative ability of CPN-CD to screen for NAFLD, patients were split into two CAP categories on TE; CAP ≥ 248 dB/m or ≥ 300 dB/m. Results: In the four cohorts, the prevalence of NAFLD ranged from 32% to 76%. Using logistic regression, the relationship of CPN-CD to CAP ≥ 248 dB/m and CAP ≥ 300 dB/m revealed C-statistic 0.80 (0.75 – 0.84) and 0.79 (0.73 – 0.85) respectively. However, at CAP ≥ 300 dB/m, CPN-CD had higher specificity (74%) and NPV (80%) compared to CAP ≥ 248 dB/m (Table). For the group who first underwent screening MRI-PDFF the yield of finding NAFLD was 2- to 6-fold higher compared to TE alone. Conclusion: CPN-CD provides fair discrimination to detect NAFLD determined on TE at CAP ≥ 300 dB/m in an external validation study conducted in four multinational cohorts. Future directions include synchronous MRI-PDFF and TE to recalibrate the score to improve specificity and then test generalizability in ulcerative colitis. Table 1. - Diagnostic accuracy of CPN-CD in detecting NAFLD in CD patients using transient elastography as reference standard CAP ≥ 248 dB/m CAP ≥ 300 dB/m Sensitivity 36% 29% Specificity 17% 74% PPV 21% 22% NPV 30% 80% CAP, controlled attenuation parameter; CD, Crohn’s disease; CPN-CD, clinical predictor tool for NAFLD in Crohn’s disease; NPV, negative predictive value; PPV, positive predictive value.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".