MétaCan
Menu
Back to cohort
Record W4388909096 · doi:10.1177/17407745231208397

Adherence to key recommendations for design and analysis of stepped-wedge cluster randomized trials: A review of trials published 2016–2022

2023· review· en· W4388909096 on OpenAlexafffund
Pascale Nevins, Mary Ryan, Kendra Davis‐Plourde, Yongdong Ouyang, Jules Antoine Pereira Macedo, Can Meng, Guangyu Tong, Xueqi Wang, Luis Ortiz‐Reyes, Agnès Caille, Fan Li, Monica Taljaard

Bibliographic record

VenueClinical Trials · 2023
Typereview
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsUniversity of OttawaOttawa Hospital
FundersNational Center for Advancing Translational SciencesPatient-Centered Outcomes Research InstituteNational Institute on AgingYale UniversityNational Institutes of HealthCanadian Institutes of Health Research
KeywordsCRTSRandomized controlled trialCovariateMedicineResearch designProtocol (science)Clinical trialCluster randomised controlled trialCluster (spacecraft)Clinical study designMedical physicsComputer scienceStatisticsData miningAlternative medicineMathematicsSurgeryPathology

Abstract

fetched live from OpenAlex

BACKGROUND/AIMS: The stepped-wedge cluster randomized trial (SW-CRT), in which clusters are randomized to a time at which they will transition to the intervention condition - rather than a trial arm - is a relatively new design. SW-CRTs have additional design and analytical considerations compared to conventional parallel arm trials. To inform future methodological development, including guidance for trialists and the selection of parameters for statistical simulation studies, we conducted a review of recently published SW-CRTs. Specific objectives were to describe (1) the types of designs used in practice, (2) adherence to key requirements for statistical analysis, and (3) practices around covariate adjustment. We also examined changes in adherence over time and by journal impact factor. METHODS: We used electronic searches to identify primary reports of SW-CRTs published 2016-2022. Two reviewers extracted information from each trial report and its protocol, if available, and resolved disagreements through discussion. RESULTS: We identified 160 eligible trials, randomizing a median (Q1-Q3) of 11 (8-18) clusters to 5 (4-7) sequences. The majority (122, 76%) were cross-sectional (almost all with continuous recruitment), 23 (14%) were closed cohorts and 15 (9%) open cohorts. Many trials had complex design features such as multiple or multivariate primary outcomes (50, 31%) or time-dependent repeated measures (27, 22%). The most common type of primary outcome was binary (51%); continuous outcomes were less common (26%). The most frequently used method of analysis was a generalized linear mixed model (112, 70%); generalized estimating equations were used less frequently (12, 8%). Among 142 trials with fewer than 40 clusters, only 9 (6%) reported using methods appropriate for a small number of clusters. Statistical analyses clearly adjusted for time effects in 119 (74%), for within-cluster correlations in 132 (83%), and for distinct between-period correlations in 13 (8%). Covariates were included in the primary analysis of the primary outcome in 82 (51%) and were most often individual-level covariates; however, clear and complete pre-specification of covariates was uncommon. Adherence to some key methodological requirements (adjusting for time effects, accounting for within-period correlation) was higher among trials published in higher versus lower impact factor journals. Substantial improvements over time were not observed although a slight improvement was observed in the proportion accounting for a distinct between-period correlation. CONCLUSIONS: Future methods development should prioritize methods for SW-CRTs with binary or time-to-event outcomes, small numbers of clusters, continuous recruitment designs, multivariate outcomes, or time-dependent repeated measures. Trialists, journal editors, and peer reviewers should be aware that SW-CRTs have additional methodological requirements over parallel arm designs including the need to account for period effects as well as complex intracluster correlations.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.977
metaresearch head score (Gemma)0.992
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Meta-epidemiology (broad), Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Meta-epidemiology (broad)
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.544
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.9770.992
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.2440.096
Bibliometrics0.0040.012
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0040.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0200.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.981
GPT teacher head0.747
Teacher spread0.234 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
DomainMethods
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations22
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueClinical TrialsSame topicMeta-analysis and systematic reviewsFrench-language works237,207