The role of pragmatism in explaining heterogeneity in meta-analyses of randomised trials: a protocol for a cross-sectional methodological review
Bibliographic record
Abstract
INTRODUCTION: There has been increasing interest in pragmatic trials methodology. As a result, tools such as the Pragmatic-Explanatory Continuum Indicator Summary-2 (PRECIS-2) are being used prospectively to help researchers design randomised controlled trials (RCTs) within the pragmatic-explanatory continuum. There may be value in applying the PRECIS-2 tool retrospectively in a systematic review setting as it could provide important information about how to pool data based on the degree of pragmatism. OBJECTIVES: To investigate the role of pragmatism as a source of heterogeneity in systematic reviews by (1) identifying systematic reviews with meta-analyses of RCTs that have moderate to high heterogeneity, (2) applying PRECIS-2 to RCTs of systematic reviews, (3) evaluating the inter-rater reliability of PRECIS-2, (4) determining how much of this heterogeneity may be explained by pragmatism. METHODS: ≥50%). Of the eligible systematic reviews, a random selection of 10 will be included for quantitative evaluation. In each systematic review, RCTs will be scored using the PRECIS-2 tool, in duplicate. Agreement between raters will be measured using the intraclass correlation coefficient. Subgroup analyses and meta-regression will be used to evaluate how much variability in the primary outcome may be due to pragmatism. DISSEMINATION: This review will be among the first to evaluate the PRECIS-2 tool in a systematic review setting. Results from this research will provide inter-rater reliability information about PRECIS-2 and may be used to provide methodological guidance when dealing with pragmatism in systematic reviews and subgroup considerations. On completion, this review will be submitted to a peer-reviewed journal for publication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Methods · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Systematic review | low |
| gpt | Metaresearch Domain: Methods · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.305 | 0.444 |
| Meta-epidemiology (narrow) | 0.006 | 0.007 |
| Meta-epidemiology (broad) | 0.011 | 0.027 |
| Bibliometrics | 0.016 | 0.017 |
| Science and technology studies | 0.004 | 0.008 |
| Scholarly communication | 0.008 | 0.008 |
| Open science | 0.007 | 0.007 |
| Research integrity | 0.014 | 0.016 |
| Insufficient payload (model declined to judge) | 0.050 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".