Meta-research on patient-reported outcomes in trial protocols and results publications suggested large outcome reporting bias
Bibliographic record
Abstract
Objectives Patient-reported outcomes (PROs) provide crucial information for evaluating health-care interventions, but previous research in specific disease areas suggested infrequent use and incomplete reporting of PROs. We examined the prevalence and characteristics of PROs in randomized clinical trial (RCT) protocols across medical fields, their reporting quality, and the consistency between PROs specified in trial protocols and subsequent reporting in trial publications. Study Design And Setting We included 237 RCT protocols approved in 2012 and 251 approved in 2016, by ethics committees in Switzerland, Germany, and Canada. We systematically searched for corresponding peer-reviewed results publications and results on trial registries. Pairs of reviewers independently extracted characteristics of RCT protocols, PROs specified in protocols and reported in corresponding results publications, and assessed the reporting quality of RCTs with a PRO as the primary outcome using the Consolidated Standards of Reporting Trials-patient-reported outcome (CONSORTs-PRO) extension. Results Out of 488 included RCT protocols, 147 (30%) did not report use of a PRO; 97 (20%) specified a PRO as the primary outcome and an additional 244 (50%) as a secondary outcome. The prevalence of PROs varied substantially across medical fields, ranging from 100% in rheumatology and psychiatry to about one-third in cardiology and anesthesiology. At 8-10 years after RCT approval, results were available for 264 of the 341 (77%) trial protocols that prespecified PROs. Forty-four percent of the published trials (115/264) reported all PROs as defined in the protocol, 21% (55/264) did not report any prespecified PROs, and 36% (94/264) reported more, fewer, or different PROs than those prespecified. These findings were consistent between trial protocols approved in 2012 and 2016. Among 63 peer-reviewed RCT publications that reported a PRO as their primary outcome, reporting quality was often inadequate, with seven of 13 CONSORT-PRO items. Conclusion Less than half of RCT protocols with planned PROs reported them as specified in corresponding published results, suggesting outcome reporting bias, and PRO reporting quality was often deficient. These limitations complicate informed decision-making between patients and health-care providers, as well as the development of evidence-based clinical practice guidelines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".