Reliability and Responsiveness of Clinical and Endoscopic Outcome Measures in Crohn’s Disease
Bibliographic record
Abstract
BACKGROUND: Regulatory guidance for Crohn's disease trials recommends coprimary efficacy end points that evaluate both symptoms and mucosal inflammation. We aimed to characterize the operating properties of commonly used disease activity assessments alone and in combination. METHODS: Endoscopic and clinical data were available for 129 participants from the Study of Biologic and Immunomodulator Naïve Patients in Crohn's Disease trial. Readers scored the Simple Endoscopic Score for Crohn's Disease and the Crohn's Disease Endoscopic Index of Severity using standardized conventions. Index reliability was determined using intraclass correlation coefficients. Index responsiveness was assessed using standardized effect sizes based upon treatment assignment. Outcomes were evaluated for optimal sensitivity to treatment effect. RESULTS: Substantial inter-rater reliability was observed when the Simple Endoscopic Score for Crohn's Disease and Crohn's Disease Endoscopic Index of Severity were used as continuous measures (intraclass correlation coefficient, 0.64; 95% confidence interval [CI], 0.50-0.73; and 0.62 95% CI, 0.36-0.77) compared with moderate reliability when dichotomized (0.46; 95% CI, 0.26-0.65; and 0.51; 95% CI, 0.00-0.78). The Simple Endoscopic Score for Crohn's Disease, Crohn's Disease Endoscopic Index of Severity, patient-reported outcome-2, and Crohn's Disease Activity Index were similarly responsive (standardized effect size, 0.43, 95% CI, 0.05-0.81; 0.38, 95% CI, 0.0-0.76; 0.53, 95% CI, 0.15-0.91). A composite outcome of Crohn's Disease Activity Index score <150 and Crohn's Disease Endoscopic Index of Severity score <6 was most sensitive to treatment effect (28.9%; 95% CI, 11.0%-46.8%; P = .003). CONCLUSION: Endoscopic indices were more reliable as continuous measures. Composite outcomes including endoscopy improved sensitivity to treatment effect.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.187 | 0.315 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".