Analysis and reporting of stratified cluster randomized trials—a systematic survey
Bibliographic record
Abstract
BACKGROUND: In order to correctly assess the effect of intervention from stratified cluster randomized trials (CRTs), it is necessary to adjust for both clustering and stratification, as failure to do so leads to misleading conclusions about the intervention effect. We have conducted a systematic survey to examine the current practices about analysis and reporting of stratified CRTs. METHOD: We used the search terms to identify the stratified CRTs from MEDLINE since the inception to July 2019. In phase 1, we screened the title and abstract for English-only studies and selected, including the main results paper of the identified protocols, for the next phase. In phase 2, we screened the full text and selected studies for data abstraction. The data abstraction form was piloted and developed using the REDCap. We abstracted data on multiple design and methodological aspects of the study including whether the primary method adjusted for both clustering and stratification, reporting of sample size, randomization, and results. RESULTS: We screened 2686 studies in the phase 1 and selected 286 studies for phase 2-among them 185 studies were selected for data abstraction. Most of the selected studies were two-arm 140/185 (76%) and parallel-group 165/185 (89%) trials. Among these 185 studies, 27 (15%) of them did not provide any sample size or power calculation, while 105 (57%) studies did not mention any method used for randomization within each stratum. Further, 43 (23%) and 150 (81%) of 185 studies did not provide the definition of all the strata, while more than 60% of the studies did not include all the stratification variable(s) in the flow chart or baseline characteristics table. More than half 114/185 (62%) of the studies did not adjust the primary method for both clustering and stratification. CONCLUSION: Stratification helps to achieve the balance among intervention groups. However, to correctly assess the intervention effect from stratified CRTs, it is important to adjust the primary analysis for both stratification and clustering. There are significant deficiencies in the reporting of methodological aspects of stratified CRTs, which require substantial improvements in several areas including definition of strata, inclusion of stratification variable(s) in the flow chart or baseline characteristics table, and reporting the stratum-specific number of clusters and individuals in the intervention groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.964 | 0.980 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.319 | 0.059 |
| Bibliometrics | 0.002 | 0.008 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".