Extending the Bayesian Adjustment for Confounding algorithm to binary treatment covariates to estimate the effect of smoking on carotid intima-media thickness: the Multi-Ethnic Study of Atherosclerosis
Bibliographic record
Abstract
We illustrate the application of the Bayesian Adjustment for Confounding (BAC) algorithm when the treatment covariate is binary. Using data from the Multi-Ethnic Study of Atherosclerosis, we estimate the effect of ever smoking on common carotid artery intimal medial thickness among adult Caucasian participants (n=1378). Our novel implementation of the BAC algorithm is performed first from an outcome model perspective and second from a treatment model perspective with both inverse probability weighting and doubly-robust estimation techniques. The BAC results are compared with the results obtained using standard model averaging and full model strategies, giving a range of adjusted estimates between 45.50 and 65.30 μm for increased common carotid artery intimal medial thickness among ever smokers. For both perspectives, we observe that BAC offers similar performance to using the fully specified outcome and/or treatment model (the full outcome model ever smoking effect is 48.61 μm; 95% CI: (0.62, 96.60)). We then redo the analyses for the African American, Hispanic, and Chinese adult participants to study the robustness of these findings with reduced sample size. For the Chinese subcohort, which corresponds to the smallest sample size (n=436), we find that, from a treatment model perspective, BAC reduces the variability of the estimates in comparison with using a full model approach. This suggests that the use of BAC in conjunction with inverse probability weighting and doubly-robust estimation can be advantageous when applied to relatively small sample sizes. This conjecture is subsequently verified on the basis of three simulated experiments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".