P546 Comparative efficacy and safety of upadacitinib versus ustekinumab as induction therapy in patients with moderately to severely active ulcerative colitis: A matching-adjusted indirect comparison
Bibliographic record
Abstract
Abstract Background Upadacitinib (UPA), an oral Janus kinase inhibitor, and ustekinumab (UST), a parenteral IL-12/23 inhibitor, are approved therapies for patients (pts) with moderately to severely active Ulcerative Colitis (UC). In the absence of head-to-head data, we conducted a placebo (PBO)-anchored matching-adjusted indirect comparison (MAIC) to compare efficacy and safety of UPA vs UST during induction. Methods Pts received UPA 45 mg oral, once daily for 8 weeks or PBO in the phase 3 studies U-ACHIEVE (NCT02819635) and U-ACCOMPLISH (NCT03653026), and UST 6 mg/kg IV at week 0 with 8-week follow up or PBO in the phase 3 study UNIFI (NCT02407236). Efficacy outcomes were stratified by previous biologic exposure (ie, no prior exposure [bio-naïve] or inadequate response, loss of response, or intolerance to biologics [bio-failed]). Baseline characteristics for age, gender, extent and duration of disease, total Inflammatory Bowel Disease Questionnaire score, total Mayo score, high-sensitivity CRP, faecal calprotectin levels, weight, and UC medication use from the UPA trials were weighted to match the UNIFI trial. Efficacy outcomes at Week 8 were clinical remission per full Mayo score (FMS; FMS ≤2 with no subscore >1), clinical response (decrease in FMS ≥3 points and ≥30%, plus a decrease in rectal bleeding score [RBS] of ≥1 or an absolute RBS of 0 or 1), endoscopic improvement (endoscopic subscore 0 or 1), and histologic-endoscopic mucosal improvement (HEMI). Safety outcomes evaluated in the overall population were adverse events (AEs) and serious AEs (SAEs), comprising the only safety assessments consistently collected across trials. Numbers needed to treat/harm (NNT/NNH) were calculated as the inverse of the difference in proportions achieving efficacy/safety outcomes between UPA and UST. Results The MAIC included 824 UPA and 625 UST pts with efficacy data and 839 UPA and 639 UST pts with safety data. A greater proportion of pts receiving UPA vs UST in bio-naïve and bio-failed groups achieved clinical remission, endoscopic improvement, and HEMI after weighting (p≤0.01, Table 1) with NNTs <10. In bio-naïve pts, differences in the proportion of pts with clinical remission, clinical response, endoscopic improvement, and HEMI for UPA vs UST were 0.172, 0.146, 0.312, and 0.255, respectively, and 0.105, 0.235, 0.135, and 0.118 for bio-failed pts. AEs and SAEs were not statistically different between UPA and UST. Conclusion During induction, greater efficacy with similar safety was observed with UPA vs UST for pts with moderately to severely active UC based on MAIC. Potential bias due to unobserved confounders may exist with MAIC methodology. Further analyses to assess longer-term outcomes are warranted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.009 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.009 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".