Reporting, handling, and interpretation of time‐varying drug treatments in observational studies using routinely collected healthcare data
Bibliographic record
Abstract
BACKGROUND: Time-varying drug treatments are common in studies using routinely collected health data (RCD) for assessing treatment effects. This study aimed to examine how these studies reported, handled, and interpreted time-varying drug treatments. METHODS: A systematic search was conducted on PubMed from 2018 to 2020. Eligible studies were those used RCD to explore drug treatment effects. We summarized the reporting characteristics and methods employed for handling time-varying treatments. Logistic regressions were performed to investigate the association between study characteristics and the reporting of time-varying treatments. RESULTS: Two hundred and fifty-six studies were included, and 225 (87.9%) studies involved time-varying treatments. Of these, 24 (10.7%) reported the proportion of time-varying treatments and 105 (46.7%) reported methods used to handle time-varying treatments. Multivariable logistic regression showed that medical studies, prespecified protocol, and involvement of methodologists were associated with a higher likelihood of reporting the methods applied to handle time-varying treatments. Among the 105 studies that reported methods, as-treated analyses were the most commonly used analysis sets, which were employed in 73.9%, 75.3% and 88.2% of studies that reported approaches for treatment discontinuation, treatment switching and treatment add-on. Among the 225 studies involved time-varying treatments, 27 (12.0%) acknowledged the potential bias introduced by treatment change, of which 14 (51.9%) suggested that potential biases may impact acceptance or rejection of the null hypothesis. CONCLUSIONS: Among observational studies using RCD, the underreporting about the presence and methods for handling time-varying treatments was largely common. The potential biases due to time-varying treatments have frequently been disregarded. Collaborative endeavors are strongly needed to enhance the prevailing practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Reporting · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Systematic review | high |
| gpt | Metaresearch Domain: Reporting · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Observational | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.287 | 0.666 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.007 |
| Bibliometrics | 0.012 | 0.019 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".