Constructing common cohorts from trials with overlapping eligibility criteria: implications for comparing effect sizes between trials
Bibliographic record
Abstract
BACKGROUND: Comparing findings from separate trials is necessary to choose among treatment options, however differences among study cohorts may impede these comparisons. PURPOSE: As a case study, to examine the overlap of study cohorts in two large randomized controlled clinical trials that assess interventions to reduce risk of major cardiovascular disease events in adults with type 2 diabetes in order to explore the feasibility of cross-trial comparisons METHODS: The Action for Health in Diabetes (Look AHEAD) and The Action to Control Cardiovascular Risk in Diabetes (ACCORD) trials enrolled 5145 and 10,251 adults with type 2 diabetes, respectively. Look AHEAD assesses the efficacy of an intensive lifestyle intervention designed to produce weight loss; ACCORD tests pharmacological therapies for control of glycemia, hyperlipidemia, and hypertension. Incidence of major cardiovascular disease events is the primary outcome for both trials. A sample was constructed to include participants from each trial who appeared to meet eligibility criteria and be appropriate candidates for the other trial's interventions. Demographic characteristics, health status, and outcomes of members and nonmembers of this constructed sample were compared. RESULTS: Nearly 80% of Look AHEAD participants were projected to be ineligible for ACCORD; ineligibility was primarily due to better glycemic control or no early history of cardiovascular disease. Approximately 30% of ACCORD participants were projected to be ineligible for Look AHEAD, often for reasons linked to poorer health. The characteristics of participants projected to be jointly eligible for both trials continued to reflect differences between trials according to factors likely linked to retention, adherence, and study outcomes. LIMITATIONS: Accurate ascertainment of cross-trial eligibility was hampered by differences between protocols. CONCLUSIONS: Despite several similarities, the Look AHEAD and ACCORD cohorts represent distinct populations. Even within the subsets of participants who appear to be eligible and appropriate candidates for trials of both modes of intervention, differences remained. Direct comparisons of results from separate trials of lifestyle and pharmacologic interventions are compromised by marked differences in enrolled cohorts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Methods · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Theoretical or conceptual | low |
| gpt | Meta-epidemiology (broad) Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | medium |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.135 | 0.219 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.017 | 0.004 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".