Statistical Treatment of Exacerbations in Therapeutic Trials of Chronic Obstructive Pulmonary Disease
Bibliographic record
Abstract
Randomized trials and a meta-analysis suggesting that inhaled corticosteroids reduce exacerbation rates in patients with chronic obstructive pulmonary disease (COPD) show major discrepancies that may be due to different approaches to data analysis. These trials used statistical techniques that were either weighted or unweighted for follow-up time, with p values and confidence intervals estimated with or without accounting for between-patient variability in exacerbation rates. We illustrate the validity of these methods using data from a cohort of 5,454 patients with COPD structured to emulate a randomized trial. The "reference" group was defined as patients with a history of exacerbations before cohort entry (n=1,137), whereas the "treated" group included an equal number (n=1,137) of patients with no prior exacerbation. Random samples of 100 and 200 subjects were selected three times from each of two groups to further illustrate the variability in the findings. Exacerbations during follow-up were identified from prescriptions for systemic antibiotics. The correct rate ratio of 0.75 estimated by the weighted approach was underestimated as 0.57 by the unweighted approach. When the weighted approach did not, however, also account for between-patient variability, the p value was greatly underestimated (e.g., rate ratio, 0.79; p=0.0007 instead of p=0.12) and confidence intervals were much narrower than after properly accounting for this variability. In conclusion, the reports from randomized trials and the meta-analysis that inhaled corticosteroids reduce COPD exacerbation rates are the result of improper statistical analysis techniques. The only two studies that used the correct statistical approach found insignificant effects with these drugs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.004 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".