Establishing a Multiple Outcome Set for Crohn’s Disease in Real-World Evidence Studies: Results from a Delphi e-Survey
Bibliographic record
Abstract
BACKGROUND: Randomized controlled trials (RCTs) provide high-quality evidence but often lack generalizability to real-world populations. Although real-world evidence (RWE) studies help to bridge this gap, retrospective design and heterogeneous outcome measures still limit their standardization in Crohn's disease (CD). Building on the recent ECCO Position Paper, this study aimed to identify the most relevant outcomes for real-world CD studies. METHODS: An international panel of inflammatory bowel disease (IBD) specialists participated in a structured two-round Delphi e-survey using the RAND/UCLA Appropriateness Method. Experts rated outcomes across eight domains, including disease activity, patient-reported outcomes, and treatment safety. Agreement was assessed using the Disagreement Index (DI), where DI > 1 indicated disagreement, and DI ≤ 1 indicated agreement or no disagreement. Weighted scoring prioritized key outcomes. RESULTS: A total of 51/85 experts (60%) completed Round 1 and 48/51 (94%) Round 2. No disagreement was observed (DI < 1) in both rounds. The highest-ranked outcomes were Abscess or Fistula (10.6%), Endoscopic Remission (10.3%), Corticosteroid-Free Clinical Remission (8.9%), Disease Progression (6.7%), and Colorectal Cancer (5.9%). The top 10 outcomes accounted for 61.5% of the weighted score. For combinations, the top four outcomes, Corticosteroid-Free Clinical Remission (16.2%), Endoscopic Remission (15.6%), Disease Progression (14.1%), and Health-Related Quality of Life (11.9%), represented 57.8% of selections. When considering the top five and top six outcomes, the cumulative proportions were 55.4% and 57.6%, respectively. CONCLUSIONS: This expert-driven Delphi study provides a standardized framework for selecting outcomes in CD RWE studies, improving consistency and comparability across future research in this field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.526 | 0.558 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.009 | 0.005 |
| Science and technology studies | 0.004 | 0.005 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.002 | 0.020 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".