MétaCan
Menu
Back to cohort

Reminding Peer Reviewers of Reporting Guideline Items to Improve Completeness in Published Articles

2023· article· en· W4379967100 on OpenAlexaff
Benjamin Speich, Erika Mann, Christof Schönenberger, Katie Mellor, Alexandra Griessbach, Paula Dhiman, Pooja Gandhi, Szimonetta Lohner, Arnav Agarwal, Ayodele Odutayo, Iratxe Puebla, Alejandra Clark, An‐Wen Chan, Michael Maia Schlüssel, Philippe Ravaud, David Moher, Matthias Briel, Isabelle Boutron, Sara Schroter, Sally Hopewell

Bibliographic record

VenueJAMA Network Open · 2023
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsOttawa HospitalUniversity of OttawaWomen's College HospitalMcMaster UniversityToronto Rehabilitation InstituteToronto General HospitalUniversity Health NetworkUniversity of TorontoImpact
FundersMagyar Tudományos AkadémiaNemzeti Kutatási Fejlesztési és Innovációs HivatalUniversität BaselSchweizerischer Nationalfonds zur Förderung der Wissenschaftlichen ForschungNational Research, Development and Innovation OfficeNational Science Foundation
KeywordsConsolidated Standards of Reporting TrialsRandomized controlled trialGuidelineMedicineRandomizationPsychological interventionProtocol (science)Family medicineClinical trialAlternative medicineNursingSurgeryInternal medicinePathology

Abstract

fetched live from OpenAlex

Importance: Numerous studies have shown that adherence to reporting guidelines is suboptimal. Objective: To evaluate whether asking peer reviewers to check if specific reporting guideline items were adequately reported would improve adherence to reporting guidelines in published articles. Design, Setting, and Participants: Two parallel-group, superiority randomized trials were performed using manuscripts submitted to 7 biomedical journals (5 from the BMJ Publishing Group and 2 from the Public Library of Science) as the unit of randomization, with peer reviewers allocated to the intervention or control group. Interventions: The first trial (CONSORT-PR) focused on manuscripts that presented randomized clinical trial (RCT) results and reported following the Consolidated Standards of Reporting Trials (CONSORT) guideline, and the second trial (SPIRIT-PR) focused on manuscripts that presented RCT protocols and reported following the Standard Protocol Items: Recommendations for Interventional Trials (SPIRIT) guideline. The CONSORT-PR trial included manuscripts that described RCT primary results (submitted July 2019 to July 2021). The SPIRIT-PR trial included manuscripts that contained RCT protocols (submitted June 2020 to May 2021). Manuscripts in both trials were randomized (1:1) to the intervention or control group; the control group received usual journal practice. In the intervention group of both trials, peer reviewers received an email from the journal that asked them to check whether the 10 most important and poorly reported CONSORT (for CONSORT-PR) or SPIRIT (for SPIRIT-PR) items were adequately reported in the manuscript. Peer reviewers and authors were not informed of the purpose of the study, and outcome assessors were blinded. Main Outcomes and Measures: The difference in the mean proportion of adequately reported 10 CONSORT or SPIRIT items between the intervention and control groups in published articles. Results: In the CONSORT-PR trial, 510 manuscripts were randomized. Of those, 243 were published (122 in the intervention group and 121 in the control group). A mean proportion of 69.3% (95% CI, 66.0%-72.7%) of the 10 CONSORT items were adequately reported in the intervention group and 66.6% (95% CI, 62.5%-70.7%) in the control group (mean difference, 2.7%; 95% CI, -2.6% to 8.0%). In the SPIRIT-PR trial, of the 244 randomized manuscripts, 178 were published (90 in the intervention group and 88 in the control group). A mean proportion of 46.1% (95% CI, 41.8%-50.4%) of the 10 SPIRIT items were adequately reported in the intervention group and 45.6% (95% CI, 41.7% to 49.4%) in the control group (mean difference, 0.5%; 95% CI, -5.2% to 6.3%). Conclusions and Relevance: These 2 randomized trials found that it was not useful to implement the tested intervention to increase reporting completeness in published articles. Other interventions should be assessed and considered in the future. Trial Registration: ClinicalTrials.gov Identifiers: NCT05820971 (CONSORT-PR) and NCT05820984 (SPIRIT-PR).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.514
metaresearch head score (Gemma)0.324
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Scholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Insufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.191
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.5140.324
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0040.001
Bibliometrics0.0000.009
Science and technology studies0.0000.000
Scholarly communication0.0030.001
Open science0.0040.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.769
GPT teacher head0.565
Teacher spread0.204 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations44
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJAMA Network OpenSame topicMeta-analysis and systematic reviewsFrench-language works237,207