MétaCan
Menu
Back to cohort
Record W4386688562 · doi:10.1186/s12874-023-02027-y

Comparing analytical strategies for balancing site-level characteristics in stepped-wedge cluster randomized trials: a simulation study

2023· article· en· W4386688562 on OpenAlexafffund
Clement Ma, Alina Lee, Darren Courtney, David Castle, Wei Wang

Bibliographic record

VenueBMC Medical Research Methodology · 2023
Typearticle
Languageen
FieldMathematics
TopicStatistical Methods and Bayesian Inference
Canadian institutionsPublic Health OntarioUniversity of TorontoCentre for Addiction and Mental Health
FundersCanadian Institutes of Health ResearchCundill Centre for Child and Youth DepressionUniversity of TorontoCentre for Addiction and Mental Health
KeywordsStatisticsSample size determinationCluster (spacecraft)MathematicsCorrelationCluster randomised controlled trialRank correlationRandomized controlled trialMedicineComputer scienceSurgery

Abstract

fetched live from OpenAlex

BACKGROUND: Stepped-wedge cluster randomized trials (SWCRTs) are a type of cluster-randomized trial in which clusters are randomized to cross-over to the active intervention sequentially at regular intervals during the study period. For SWCRTs, sequential imbalances of cluster-level characteristics across the random sequence of clusters may lead to biased estimation. Our study aims to examine the effects of balancing cluster-level characteristics in SWCRTs. METHODS: To quantify the level of cluster-level imbalance, a novel imbalance index was developed based on the Spearman correlation and rank regression of the cluster-level characteristic with the cross-over timepoints. A simulation study was conducted to assess the impact of sequential cluster-level imbalances across different scenarios varying the: number of sites (clusters), sample size, number of cross-over timepoints, site-level intra-cluster correlation coefficient (ICC), and effect sizes. SWCRTs assumed either an immediate "constant" treatment effect, or a gradual "learning" treatment effect which increases over time after crossing over to the active intervention. Key performance metrics included the relative root mean square error (RRMSE) and relative mean bias. RESULTS: Fully-balanced designs almost always had the highest efficiency, as measured by the RRMSE, regardless of the number of sites, ICC, effect size, or sample sizes at each time for SWCRTs with learning effect. A consistent decreasing trend of efficiency was observed by increasing RRMSE as imbalance increased. For example, for a 12-site study with 20 participants per site/timepoint and ICC of 0.10, between the most balanced and least balanced designs, the RRMSE efficiency loss ranged from 52.5% to 191.9%. In addition, the RRMSE was decreased for larger sample sizes, larger number of sites, smaller ICC, and larger effect sizes. The impact of pre-balancing diminished when there was no learning effect. CONCLUSION: The impact of pre-balancing on preventing efficiency loss was easily observed when there was a learning effect. This suggests benefit of pre-balancing with respect to impacting factors of treatment effects.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.154
metaresearch head score (Gemma)0.313
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.846
Threshold uncertainty score0.817

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1540.313
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0020.002
Science and technology studies0.0010.002
Scholarly communication0.0020.003
Open science0.0030.003
Research integrity0.0030.003
Insufficient payload (model declined to judge)0.0040.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.872
GPT teacher head0.664
Teacher spread0.208 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueBMC Medical Research MethodologySame topicStatistical Methods and Bayesian InferenceFrench-language works237,207