The evolving context of reporting randomised trials. A scoping review of comments on SPIRIT 2013 and CONSORT 2010
Bibliographic record
Abstract
Objective: To identify, summarise, and analyse comments on the core reporting guidelines for protocols of randomised trials (SPIRIT 2013) and for completed trials (CONSORT 2010), with special emphasis on suggestions for guideline modifications.Methods: We included documents written in English and published after 2010 that explicitly commented on SPIRIT 2013 or CONSORT 2010. We searched four bibliographic databases (Embase and MEDLINE to June 2022; Web of Science and Google Scholar to April 2022) and other sources (e.g., the EQUATOR Network web-site, the BMC Blog Network, and the BMJ rapid response section). Two authors independently assessed documents for eligibility and extracted data on basic characteristics and the wording of the main comments. We categorized comments as ‘suggestion for modification to the wording of an existing guideline item’, ‘suggestion for a new item’, or ‘reflections on challenges or strengths’. We provided a summary and examples of the proposed suggestions and categorised comments into those that were directly linked to empirical investigations, were continuations of previous methodological discussions, or reflected new methodological developments.Results: We assessed full text of 2320 potentially eligible documents and included 93 documents with 114 comments. In total, 37 comments suggested modifications to existing guideline items. The participant flow section of CONSORT 2010 received the most comments (eight comments made different suggestions, e.g., one comment suggested to add numbers on non-randomised screened participants). There were 46 comments suggesting new items. Multiple suggestions were related to trial interventions (eight comments made different suggestions, e.g., one comment suggested to add content on co-interventions), blinding (six comments suggested to add content on risk of unblinding), statistical methods (five comments made different suggestions, e.g., one comment suggested to add content on blinding of statisticians), and participant flow (seven comments made different suggestions, e.g., three comments suggested to add content on missing data). Half (53%) of the suggestions were directly linked to empirical investigations. Six (7%) suggestions were continuations of previous methodological discussions, and five (6%) suggestions reflected new methodological developments related to conflicts of interest and funding, data sharing, and patient and public involvement.Conclusion: The issues raised provide context to authors, peer reviewers, editors and readers of trials using SPIRIT 2013 and CONSORT 2010 and inform the planned updates of the core guidelines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.751 | 0.914 |
| Meta-epidemiology (narrow) | 0.003 | 0.006 |
| Meta-epidemiology (broad) | 0.007 | 0.008 |
| Bibliometrics | 0.032 | 0.030 |
| Science and technology studies | 0.008 | 0.020 |
| Scholarly communication | 0.019 | 0.026 |
| Open science | 0.010 | 0.018 |
| Research integrity | 0.022 | 0.026 |
| Insufficient payload (model declined to judge) | 0.005 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".