Reporting, handling, and interpretation of time‐varying drug treatments in observational studies using routinely collected healthcare data
Bibliographic record
Abstract
BACKGROUND: Time-varying drug treatments are common in studies using routinely collected health data (RCD) for assessing treatment effects. This study aimed to examine how these studies reported, handled, and interpreted time-varying drug treatments. METHODS: A systematic search was conducted on PubMed from 2018 to 2020. Eligible studies were those used RCD to explore drug treatment effects. We summarized the reporting characteristics and methods employed for handling time-varying treatments. Logistic regressions were performed to investigate the association between study characteristics and the reporting of time-varying treatments. RESULTS: Two hundred and fifty-six studies were included, and 225 (87.9%) studies involved time-varying treatments. Of these, 24 (10.7%) reported the proportion of time-varying treatments and 105 (46.7%) reported methods used to handle time-varying treatments. Multivariable logistic regression showed that medical studies, prespecified protocol, and involvement of methodologists were associated with a higher likelihood of reporting the methods applied to handle time-varying treatments. Among the 105 studies that reported methods, as-treated analyses were the most commonly used analysis sets, which were employed in 73.9%, 75.3% and 88.2% of studies that reported approaches for treatment discontinuation, treatment switching and treatment add-on. Among the 225 studies involved time-varying treatments, 27 (12.0%) acknowledged the potential bias introduced by treatment change, of which 14 (51.9%) suggested that potential biases may impact acceptance or rejection of the null hypothesis. CONCLUSIONS: Among observational studies using RCD, the underreporting about the presence and methods for handling time-varying treatments was largely common. The potential biases due to time-varying treatments have frequently been disregarded. Collaborative endeavors are strongly needed to enhance the prevailing practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Reporting · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
| gpt | Metaresearch Domain: Reporting · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Observational | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.048 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".