P176 Expert consensus on acquisition and reporting of intestinal ultrasonography activity in Crohn’s disease. A prospective inter-rater agreement study
Bibliographic record
Abstract
Abstract Background Intestinal ultrasonography (IUS) is a promising cross-sectional imaging modality used to assess transmural disease and complications in Crohn’s disease (CD). Although recently positioned as a first-line modality for evaluation as per ECCO guidelines, standard measurements, reproducibility and nomenclature have not yet been clearly established. The aim of this study was to evaluate the inter-rater agreement for parameters identified as important by experts through Delphi consensus. Methods IUS parameters demonstrating inflammatory activity were systematically reviewed in the literature and presented to IUS experts. Individual parameters were selected by a blinded Delphi consensus panel to establish relative contribution to inflammatory activity in CD. Weighted grading of each parameter was further established by expert consensus. Image acquisition for optimal measurement was established by consensus. Two phases for evaluating inter-rater variability were undertaken. Phase 1: blind review by 8 readers of 20 de-identified CD cases. Cases with poor agreement were reviewed to clarify discrepancy and improve agreement. Phase 2: an additional 30 de-identified CD cases blindly were reviewed by 12 independent expert readers. Inter-rater agreement was evaluated for all 4 key parameters. Statistics were performed using Stata 16. Bowel wall thickness (BWT) was assessed using intraclass correlation coefficient (ICC) and the ordinal parameters using weighted Cohens Kappa. Results The Delphi process reduced 12 activity parameters to 4 key contributors including BWT, color Doppler signal (CDI), inflammatory fat and bowel wall echostratification (Figure 1). BWT was regarded as pathologic if the average of 4 measurements were > 3 mm for the small and large bowel, and grades of the additional parameters established (Table 1). Bowel wall thickness was comprised of 2 measurements in cross section and 2 in longitudinal orientation (Figure 2). Interobserver agreement was almost perfect for BWT: ICC=0.91 (95% CI 0.83 to 0.96) p = 0.001, while there was moderate agreement for CDI κ=0.60 (95% CI 0.48–0.72) p = 0.001. Agreement for inflammatory fat detection was also moderate with κ= 0.50 (95% CI 0.33–0.66) p = 0.001, while stratification was fair κ= 0.39 (95% CI 0.26–0.53) p = 0.001. Conclusion This expert consensus-based IUS activity score clearly establishes the reproducibility of this standardised approach to measure inflammatory activity in patients with CD. Using our method, BWT which is known as the most important parameter, is highly reproducible with CDI and inflammatory fat demonstrating moderate reproducibility. This score may provide the foundation for the future incorporation of IUS in research studies and clinical trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.284 | 0.341 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.007 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".