DOP12 Validation of radiomics features on MR enterography characterizing inflammation and fibrosis in stricturing Crohn’s disease
Bibliographic record
Abstract
Abstract Background MR enterography (MRE) accurately detects Crohn’s disease (CD) strictures, yet its ability to differentiate inflammatory from fibrostenotic components within a CD stricture is limited. Artificial intelligence in cross sectional imaging, termed radiomics, is a quantitative image extraction analysis technology creating an opportunity to enhance characterization of strictures on routine MRE exams. We present a study on machine-reader evaluation of MRE to distinguish inflammation and fibrosis in CD strictures via quantitative radiomic features and compare radiomics performance to central radiologist scoring of MRE. Methods In this retrospective single center study 51 patients (n=34 for discovery; n=17 for validation) had confirmed stricturing CD (using CONSTRICT criteria) on MRE. Surgical histopathology scoring of specimens within 15 weeks of MRE exam (range 0-100, scores ≥70 =severe) was used as the reference standard for both inflammation and fibrosis. An expert abdominal radiologist blinded to clinical and histopathologic results provided a global visual analog scale (VAS, 0-100) assessment of stricture inflammation and fibrosis. 2164 3D radiomic features were extracted from the stricture regions on MRE, from which the most relevant feature subsets were identified via cross-validated machine learning analysis in the discovery cohort for differentiating between severe vs mild inflammation and fibrosis. Radiomic features and VAS scores were evaluated against pathology-defined inflammation and fibrosis in the validation cohort. Results Clinical variables including sex, age, Montreal classification and stricture type across discovery and validation groups can be found in Table 1. The median time from MRE to surgical resection was 7.1 90-15) weeks. 43% of strictures in the overall cohort were classified as severe for inflammation and 43% had severe fibrosis. Two distinct sets of radiomic features capturing textural heterogeneity (patterns, local entropy) within strictures were significantly associated with severe inflammation or severe fibrosis (p<0.01). For inflammation, AUC for discovery and validation were 0.69 and 0.67, respectively (Figure 1). For fibrosis, AUC for discovery and validation were 0.83 and 0.77, respectively (Figure 2). The radiologist VAS had an AUC of 0.71 for identifying inflammation and AUC 0.46 for identifying fibrosis. Combining radiomic features and radiologist VAS had no significant impact on predictor performance. Conclusion Radiomic analysis may support the identification of fibrosis, but not inflammation in stricturing CD compared to radiological visual assessment. This tool may offer a novel way to stratify patients for future anti-fibrotic therapies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".