Identifying heterogeneity of treatment effect for antibiotic duration in bloodstream infection: an exploratory post-hoc analysis of the BALANCE randomised clinical trial
Bibliographic record
Abstract
Background: bacterial bloodstream infections (BSI). However, there may be patient subgroups who benefit from longer durations. We aimed to evaluate if bedside clinical decision rules could identify these subgroups. Methods: In this post-hoc analysis of the multicentre, randomised BALANCE trial (October 17, 2014-May 5, 2023), we applied three clinical decision rules to investigate heterogeneity of treatment effect in 7-day vs 14-day antibiotic durations on 90-day all-cause mortality. We used the rules to categorize patients in BALANCE into different risk groups and calculated the unadjusted absolute risk difference (RD) for 90-day mortality in patients receiving 7- vs 14-day antibiotics within each risk group. Statistical significance was tested using an interaction test. The BALANCE trial is registered with ClinicalTrials.gov (NCT03005145). Findings: 3581 patients were included. All three rules predicted mortality risk, but none identified statistically significant effect modification: (a) static rule (low-risk: RD -0.58, 95% CI -8.91 to 7.73; moderate-risk: RD -.01, 95% CI -3.86 to 1.83; high-risk: RD -2.65, 95% CI -7.12 to 1.81; p = 0.74); (b) dynamic rule (met rule on day 7: RD -2.18, 95% CI -4.81 to 0.45; did not meet rule: RD 1.75, 95% CI -3.89 to 7.40; p = 0.16); and (c) early clinical failure criteria (score<2: RD -2.38, 95% CI -5.0 to 0.23; score ≥2: RD -0.65, 95% CI -5.06 to 3.77; p = 0.24). Results were consistent across sensitivity analyses including imputation for missing data and restricting analyses to gram-negative BSI. Interpretation: BSI. Future research could explore data-driven machine-learning approaches to identify comprehensive combinations of patient characteristics that may guide individualised duration of antibiotic therapy. Funding: The BALANCE trial was funded by the Canadian Institutes of Health Research, Health Research Council of New Zealand, Australian National Medical Research Council, Physicians Services Incorporated Ontario and Ontario Ministry of Health and Long-term Care Innovation Fund. SWXO conducted this study as part of his PhD studies, with funding from: the Emerging & Pandemic Infections Consortium (University of Toronto, Canada); Connaught International Scholarship (University of Toronto, Canada); the Queen Elizabeth II Graduate Scholarship in Science and Technology (QEII-GSST; Government of Ontario, Canada); and the Melbourne Research Scholarship (University of Melbourne, Australia). VML is supported by Clinical Research Scholar-Junior 2 program (FRQ-S).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".