Bias–variance trade‐off in pharmacoepidemiological studies using physician‐preference‐based instrumental variables: a simulation study
Bibliographic record
Abstract
PURPOSE: Instrumental variables (IV) methodology removes bias due to unobserved confounding by replacing in the analysis the treatment with another variable--the instrument--that is well correlated with the treatment and independent of confounders. Recently, physician drug preference, operationalized as the treatment prescribed to the previous patient of the same physician, was proposed as an instrument in database studies comparing two competing drugs. We assessed, in simulations, how the performance of the IV estimates depends on the strength of this instrument. METHODS: The 'physician preference' instrument correlates well with the treatment only if physician preferences affect the treatment received by a large fraction of patients. Yet, often there is a subgroup of patients whose treatment cannot be affected by physician's preferences. The larger this subgroup is, the weaker the instrument. We investigated the impact of weakening this instrument on the performance of IV estimates, by comparing risk difference estimates from the conventional and IV analyses in the presence of an unobserved confounder, for both continuous and binary outcomes. RESULTS: The IV estimates were uniformly less biased than the conventional estimates, but had higher variance. Accordingly, the bias-variance trade-off favors the IV estimates only when physician' preference is a strong instrument. Still, the coverage rate of the 95%CI for the IV estimates was very close to the nominal 95%, while it was consistently lower for the conventional estimates. CONCLUSION: Researchers should consider using the 'physician preference' instrument when comparing two competing drugs, but should be aware of the underlying assumptions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".