S762 Evaluation of Clinical Variables, Radiological Visual Analog Scoring, and Radiomics Features on MR Enterography for Characterizing Severe Inflammation and Fibrosis in Stricturing Crohn’s Disease
Bibliographic record
Abstract
Introduction: Current non-invasive cross-sectional imaging modalities such as MR enterography (MRE) offer excellent diagnostic accuracy of Crohn’s disease (CD) strictures, but cannot accurately determine the extent of stricture fibrosis and inflammation. Radiomics, a quantitative image extraction analysis technology, may offer a solution. We present initial results for a machine-reader evaluation of severe inflammation and fibrosis in CD strictures via quantitative radiomic features and expert radiologist scoring of MRE. Methods: In this retrospective, single center, IRB-approved study, 51 patients (n=34 for discovery; n=17 for hold-out validation) had confirmed stricturing CD on MRE and histopathology from surgery within 15 weeks of MRE. Histopathological Stenosis Therapy & Research (STAR) scoring of specimens (range 0-100, scores ≥50 =severe) was the reference standard for both inflammation and fibrosis. An expert radiologist coordinated with the scoring pathologist to annotate the resected strictures on MRE and provide a global visual analog score (VAS, 0-100) assessment of inflammation and chronic non-inflammatory findings (fibrosis). 1852 3D radiomic features were extracted from the stricture regions on MRE, from which the most relevant feature subsets were identified via cross-validated machine learning analysis in the discovery cohort for differentiating between severe vs less severe inflammation and fibrosis. Radiomic features and VAS scores were evaluated against pathology-defined severe inflammation and fibrosis in the validation cohort via ROC analysis. Results: Two distinct sets of radiomic features capturing textural heterogeneity (patterns, local entropy) within strictures were significantly associated (p< 0.01) with severe inflammation and severe fibrosis; across both discovery (AUC=0.66, 0.76) and hold-out validation (AUCs =0.71,0.83) (Figure). Radiological VAS had an AUC=0.68 for identifying severe inflammation and AUC =0.47 for severe fibrosis. Combining radiomic features and VAS had no significant impact on predictor performance. Clinical variables including sex, age, Montreal classification and stricture type were not significantly associated severe inflammation or fibrosis, across discovery and validation groups (Table). Conclusion: Radiomic analysis shows improved performance in identifying severe inflammation and severe fibrosis in CD strictures on MRE compared to radiological visual assessment scoring and clinical variables.Figure 1.: Top-ranked radiomics features are distinctively associated with severe inflammation (top row, pattern-based) and severe fibrosis (bottom row, wavelets) on MRE. Also shown are radiological VAS for severe inflammation and severe fibrosis. Table 1. - Demographics and baseline clinical features of the cohort, segregating discovery and hold-out validation radiomic cohorts MRE Overall (N=51) Fibrosis Discovery Group (N=34) Fibrosis Validation Group (N=17) Inflammation Discovery Group (N=34) Inflammation Validation Group (N=17) Factor N Statistics N Statistics N Statistics P-value N Statistics N Statistics P-value Male Sex, n (%) 51 26 (51) 34 16 (47) 17 10 (57) 0.43 a 34 15 (44) 17 11 (65) 0.17 a Diagnosis age of IBD, median (range), yrs 51 21 (4-90) 34 24 (10-90) 17 20 (2-62) 0.88 c 34 20.5 (5-67) 17 25 (4-90) 0.79 c Diagnosis age of Stricture, median (range), yrs 51 32 (11-90) 34 30.5 (19-90) 17 33 (11-69) 0.78 c 34 29.5 (11-71) 17 35 (20-90) 0.24 c Age at MRE, median (range), yrs 51 34 (18-91) 34 33 (19-91) 17 36 (18-69) 0.83 c 34 31 (18-71) 17 37 (22-91) 0.21 c Duration between IBD/stricture dx, median (range), years 51 8 (0-30) 34 6.5 (0-30) 17 10 (0-26) 0.62 c 34 7.5 (0-30) 17 8 (0-21) 0.8 c Duration between Stricture dx/Surgery, median (range), months 51 9 (0-145) 51 10 (0-145) 17 6 (0-121) 0.82 c 34 5.5 (0-145) 17 19 (0-78) 0.24 c Duration between MRE and resection, median (range), weeks 51 7.1 (0-15) 34 7.35 (0-13) 17 7 (0.9-15) 0.36 c 34 7.9 (0.1-15) 17 7.1 (0-14.9) 0.93 c Obstructive Symptoms at time of imaging, n (%) 51 42 (82) 34 28 (82) 17 14 (82) 1 b 34 27 (79) 17 15 (88) 0.7 b CD Montreal Classification, n (%) 51 34 17 0.64 a 34 17 0.14 a B2 (Stricturing) 23 (45) 14 (41) 9 (53) 17 (50) 6 (35) B2p (Stricturing with perianal disease) 15 (29) 11 (32) 4 (23) 11 (32) 4 (23) B3 (Fistulizing) 6 (12) 5 (15) 1 (6) 4 (12) 2 (12) B3p (Fistulizing with perianal disease) 7 (14) 4 (12) 3 (18) 2 (6) 5 (29) History of extraintestinal manifestations, n (%) 51 32 (63) 34 22 (65) 17 10 (59) 0.68 a 34 22 (65) 17 10 (59) 0.68 a Ileocecal resection prior to current stricture, n (%) 51 25 (49) 34 17 (50) 17 8 (47) 0.84 a 34 16 (47) 17 9 (53) 0.69 a Number of resections, median (range) 25 2 (1-5) 16 2 (1-5) 8 2 (1-4) 0.88 c 16 2 (1-4) 2 (1-5) 0.94 c Type of stricture, n (%) 51 34 17 1 a 34 17 1 a Naïve 27 (53) 18 (53) 9 (53) 18 (53) 9 (53) Anastomotic 24 (47) 16 (47) 8 (47) 16 (47) 8 (47) Medications for IBD < 8 weeks from imaging, n (%) 51 34 17 34 17 5-aminosalicylic-acid, oral or rectal 10 (20) 8 (24) 2 (12) 0.46 b 6 (18) 4 (24) 0.71 b Steroid, systematic 19 (37) 11 (32) 8 (47) 0.31 a 12 (35) 7 (41) 0.68 a Steroid, rectal or Budesonide 12 (24) 9 (27) 3 (18) 0.73 b 8 (24) 4 (24) 1 b Mercaptopurine or Azathioprine 12 (24) 7 (21) 5 (29) 0.5 b 10 (29) 2 (12) 0.29 b Methotrexate 2 (4) 2 (6) 0 (0) 0.55 b 1 (3) 1 (6) 1 b Certolizumab 2 (4) 2 (6) 0 (0) 0.55 b 1 (3) 1 (6) 1 b Adalimumab 13 (25) 10 (29) 3 (18) 0.5 b 8 (24) 5 (29) 0.74 b Infliximab 5 (10) 3 (9) 2 (12) 1 b 4 (12) 1 (6) 0.65 b Vedolizumab 5 (10) 4 (12) 1 (6) 0.65 b 2 (6) 3 (18) 0.32 b None 6 (12) 4 (12) 1 (6) 0.65 b 4 (12) 1 (6) 0.65 b Global Assessments by Radiologist Global Stricture Severity, median (range), 0-100 51 60 (20-100) 34 50 (20-100) 17 60 (20-100) 0.4 c 34 60 (20-100) 17 40 (20-95) 0.44 c Global Inflammation Severity, median (range), 0-100 51 40 (15-95) 34 40 (15-85) 17 50 (15-95) 0.27 c 34 40 (15-85) 17 40 (15-95) 0.7 c Global Chronic non-inflammatory changes Severity, median (range), 0-100 51 30 (5-80) 34 30 (5-80) 17 40 (5-80) 0.16 c 34 30 (5-80) 17 30 (5-70) 0.89 c Global Assessments by Pathologist Severity of inflammation, median (range), 0-100 51 66 (2-100) 34 67 (2-100) 17 64 (10-94) 0.52 c 34 61.5 (2-100) 17 67 (10-100) 0.73 c Severity of fibrosis, median (range), 0-100 51 60 (5-94) 34 60 (5-94) 17 55 (10-88) 0.93 c 34 57.5 (5-94) 17 63 (10-90) 0.93 c aChi-Square testbFisher exact testcMann Whitney U test.MRE: magnetic resonance enterography; N: Number; IBD: inflammatory bowel disease; dx: diagnosis; CD: Crohn’s disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".