829 Appropriateness of Treatment Options in Mild to Moderate Ulcerative Colitis: A RAND Appropriateness Panel
Bibliographic record
Abstract
INTRODUCTION: Treatment options for patients with moderate-severe ulcerative colitis (UC) are increasing. The appropriateness of their use in mild-moderate UC is less clear, as is the role of surgical intervention. Our aim was to define the appropriateness of therapeutic options in a range of clinical scenarios in mild-moderate UC METHODS: We applied the RAND/UCLA Appropriateness Method to rate appropriateness of treatment choices in patients with mild-moderate UC; severe disease (likely to need hospitalization or surgery within 2 weeks) was excluded. A literature review was presented to the BRIDGe group, a panel of 13 IBD specialists, as well as 1 external IBD expert and 2 IBD surgical experts. 12 chapters were constructed including failed 5-ASA, steroid or thiopurine therapy, and primary/secondary non-response to biologics or tofacitinib. Scenarios were also divided by disease extent (proctitis vs more extensive) and activity (mild vs moderate) creating 260 scenarios in all. Panellists used a modified Delphi method to anonymously rate the appropriateness of a therapy on a 1-9 scale (1-3 not appropriate, 4-6 uncertain, 7-9 appropriate). Disagreement was assessed using a validated index (DI) and was defined as a DI > 1. The panellists then convened in person to discuss areas of disagreement, followed by a second round of rating. Scenarios rated a median of 7 or higher were deemed appropriate and those rated 3 or lower were deemed inappropriate. Scenarios rated 4-6 were considered uncertain, as were any with a DI > 1 RESULTS: Appropriateness was agreed in 131 and inappropriateness in 47 scenarios. Greatest disagreement was around the use of thiopurines, but no scenario had a DI > 1. There were 82 scenarios were rated uncertain. Biologics and tofacitinib were more likely to be rated appropriate for treatment resistance and greater disease activity; disease extent less likely to influence decision making. Surgery was felt to be appropriate in 8 scenarios, generally in more extensive and active disease that had failed to respond to medical therapy. Surgery was felt to be inappropriate in 16 scenarios and was uncertain in 24 CONCLUSION: Although there was little disagreement amongst an expert panel of IBD physicians and surgeons, there remained uncertainty about the use of medical and surgical therapy in nearly one third of situations in patients with mild-moderate UC. These results could be used to guide the design of clinical trials for these patient groups in order to guide best practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".