Screen Failures and Causes in Inflammatory Bowel Disease Randomized Controlled Trials: A Study of 16 913 Screened Patients
Bibliographic record
Abstract
INTRODUCTION: While recruitment rates in inflammatory bowel disease (IBD) trials are continuously decreasing, the underlying reasons are likely multifactorial but remain poorly defined. Screen failure (SF) proportions and causes have not been extensively explored in IBD. AIM: We assessed SF proportions and underlying SF reasons in IBD phase 2 and 3 clinical trials. METHODS: We analyzed SF-related data from 17 randomized controlled phase 2 or 3 IBD trials. Twelve trials were in ulcerative colitis (UC) and 5 trials were in Crohn's disease (CD) operated by a single contract research organization, IQVIA. Differences between patient groups were tested for significance by Mann-Whitney and Fisher's tests when appropriate. RESULTS: We analyzed a total of 11 161 patients with UC and 5752 patients with CD. The mean SF proportion was 0.43 per trial in UC. The primary reason for SFs in UC was not meeting the overall (modified) Mayo score inclusion threshold and/or the endoscopic subscore of at least 2 (33.8% of all SF). In CD clinical trials, the mean SF proportion was at 0.53. The primary cause for SFs was not meeting the CDAI eligibility criteria (23.1% of all SFs). SF proportions were significantly higher in CD versus UC trials (P = .027). Clostridium difficile or any other intestinal infection and not meeting tuberculosis screening criteria were other major reasons for SFs both in UC and CD. CONCLUSION: High SF proportion in IBD clinical trials, particularly for CD studies, pose obstacles to patient recruitment. While underlying causes are diverse, arbitrarily defined clinical and/or endoscopic eligibility criteria remain the major limiting factors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.099 | 0.147 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.007 | 0.010 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".