A160 PREVALENCE OF OUTCOME SWITCHING AMONG PUBLISHED PHASE 3 INTERVENTIONAL TRIALS FOR INFLAMMATORY BOWEL DISEASE THERAPEUTICS
Bibliographic record
Abstract
Abstract Background Outcome switching is a well-described form of inconsistent reporting in randomized clinical trials (RCTs), wherein pre-specified primary and/or secondary outcomes are changed between trial registration and the publication of results without explanation. This is of particular concern, as the selective publication of results that are favorable will insert bias into the trial’s results and may cast doubt on the veracity of its findings. While it has been investigated in other disciplines, the prevalence of outcome switching has yet to be described among RCTs for inflammatory bowel disease (IBD). Aims To determine the prevalence of correctly reported pre-specified primary and secondary outcomes in published phase 3 interventional RCTs for IBD. Methods We identified all phase 3 interventional trials for IBD with published results using clinicaltrials.gov. We included all results with an associated publication that detailed the results of the trial. We excluded registrations if: only an abstract of the results was available; trial results were only published as a pooled analysis; multiple trial segments were reported collectively; or a publication of the results could not be identified through clinicaltrials.gov or a custom search. Two reviewers extracted all pre-specified primary and secondary outcomes for each trial using the clinical trial registration page that was dated before the commencement of the trial. These outcomes were compared to the outcomes reported in the corresponding journal articles. Any discrepancies were noted, and additional outcomes were extracted. Results We identified a total of 88 phase 3 interventional RCTs for IBD, of which 57 were matched to independent publications of their results. All trials pre-specified a primary outcome, and 50 (87.7%) pre-specified secondary outcomes. 10 (17.5%) of trials did not report some or all primary outcomes, and 19 (33.3%) trials had a change or alteration to the primary outcome. Of the trials that pre-specified secondary outcomes, 16 (28.1%) did not report all pre-specified secondary outcomes. 49 (86.0%) trials added 6 (IQR: 2–8) unspecified secondary outcomes on average. Conclusions Many phase 3 interventional RCTs in IBD either did not report some or all primary outcomes, or altered the primary outcome. Trials routinely reported additional outcomes that were not pre-specified and failed to note that they were added post hoc. Based on these results, we recommend improvements in the reporting of pre-specified outcomes and higher fidelity in order to maintain confidence in trial results. Funding Agencies None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.423 | 0.775 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.007 |
| Bibliometrics | 0.027 | 0.023 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.008 | 0.007 |
| Open science | 0.004 | 0.008 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".