NO UNMEASURED CONFOUNDING: KNOWN UNKNOWNS OR… NOT?
Bibliographic record
Abstract
A large majority of methods described in the causal literature rely on the assumption of no unmeasured confounding (NUC). When estimating treatment effects, the NUC assumption requires the measurement of all variables related to both treatment exposure and the outcome(s) of interest. In this short letter, we discuss 2 approaches which one might think could provide validation of the NUC assumption and show that neither is appropriate for this purpose. We close with a reminder to readers of the optimistic view of the complexity of data and how correlation between observed variables and unmeasured ones can reduce any bias associated with unmeasured information. We begin with some notation. Suppose we are interested in estimating an average treatment effect with a regression-based framework. Let |$Y$| denote a continuous outcome, |$X$| a measured confounder (possibly a vector), |$U$| an unmeasured confounder, and |$A$| a binary treatment. We assume the structural model>? where |${\mathrm{\varepsilon}}_i$| is a mean zero error term (see Web Appendix 1, available at https://doi.org/10.1093/aje/kwad133, for further details on the structural model). The function |$\mathrm{\gamma} \left(A,X;\mathrm{\psi} \right)$| represents the treatment effect, with |$\mathrm{\gamma} \left(A=0,X;\mathrm{\psi} \right)=0$| and |$\mathrm{\psi}$| the treatment effect estimand of interest. While |$\mathrm{\gamma} \left(A,X;\mathrm{\psi} \right)$|allows for the estimation of conditional treatment effects (i.e., conditional on |$X$|), if |$\mathrm{\gamma} \left(A,X;\mathrm{\psi} \right)$| = |$\mathrm{\gamma} \left(A;\mathrm{\psi} \right)$| then we assume an unconditional treatment effect. The function |$f({X},{U};\mathrm{\beta} )$| represents the effect of confounders in the absence of treatment (i.e., when |$A=0$|). In this letter, with some minor loss of generality, we assume that the effect of treatment, |$A$|, may be modified by |$X$| but not |$U$|. For simplicity, we assume linear relationships throughout numerical examples; however, the theory and conclusions are not constrained to that case. Assuming the structural model above permits the counterfactual formulation |${Y}_i(0)\!={Y}_i(a)-\mathrm{\gamma} \left(a,{x}_i;\mathrm{\psi} \right)$| = |$f({x}_i,{u}_i;\mathrm{\beta} )+{\mathrm{\varepsilon}}_i$|, where |${Y}_i(a)$| is the counterfactual outcome if individual |$i$| had received (potentially counter to fact) treatment |$a$|. Letting the vector of all confounders be denoted by |$\mathbf{L}=\left(X,U\right)$|, the NUC assumption entails conditional independence between the counterfactual |$Y(A)$| and treatment |$A$|, given the confounders |$\mathbf{L}$|, that is, as |$Y(A)\perp A\mid \mathbf{L}$|. Suppose, then, we had a candidate value of the true treatment effect, |$\mathrm{\psi}$|, which we denote by |${\mathrm{\psi}}^{\dagger }$|, and posited structural model (which we shall assume is linear for ease of notation). Following Hernán and Robins (1) (see section 14.5), define |$H({\mathrm{\psi}}^{\dagger})=Y-\mathrm{\gamma} \left(A,X;{\mathrm{\psi}}^{\dagger}\right)$| as a surrogate for the counterfactual |$Y(0)$| based on an assumed value for the treatment effect |${\mathrm{\psi}}^{\dagger }$|. Then |$Y(0)\perp A\mid \mathbf{L} \;$|—that is, the NUC assumption—implies that the assignment of treatment does not depend on how an individual might react to treatment, or in its absence. Therefore 1) |${\zeta}_1\!=\!0$| in |$E\left[H\left(\mathrm{\psi} \right)|A,\mathbf{L};\zeta \right]={\zeta}_0+{\zeta}_1A+{\zeta}_2^{\top}\mathbf{L}$| and 2) |${\mathrm{\alpha}}_1\!=\!0$| in |$\mathrm{logit}(\Pr [A=1|H\left(\mathrm{\psi} \right)\!,\mathbf{L};\mathrm{\alpha} ])={\mathrm{\alpha}}_0+{\mathrm{\alpha}}_1H\left(\mathrm{\psi} \right)+{\alpha}_2^\top \mathbf{L}$|when |$\mathrm{\psi}$| is the true causal treatment effect. It is tempting to use these facts as a means of testing the NUC assumption. However, in Web Appendix 2 (data generation according to Web Figure 1, yielding results presented in Web Table 1), we show straightforward and plausible examples wherein using |$H({\mathrm{\psi}}^{\dagger})$| results in |${\zeta}_1=0$| in the first case above (point 1) and |${\mathrm{\alpha}}_1=0$| in the second (point 2), even when |${\mathrm{\psi}}^{\dagger }$| is a biased estimate of the true |$\mathrm{\psi}$| due to the exclusion of an unmeasured confounder (i.e., when |${\mathrm{\psi}}^{\dagger }$| is determined from a model with |$\mathbf{L}=X$| rather than |$\mathbf{L}=\left(X,U\right)$|). An alternative idea, based on model-fitting procedures, might also be considered for assessing whether NUC holds. Building on ideas from a comment by Robins and Rotnitzky (2), Wallace et al. (3) demonstrated that double-robustness could be leveraged to diagnose model mispecification. The approach stems from the observation that, under double robustness, if one of the models needed for estimation is correct, estimators are consistent even when a second model is misspecified. Consider, for example, a setting in which either the outcome or the treatment (propensity score) model must be correctly specified. The analyst could compare treatment effect estimates resulting from a doubly robust procedure using their best posited outcome model and several candidate propensity score models, including models known or strongly suspected to be misspecified, such as a null model. If the resulting treatment effect estimates are stable—that is, given the fixed optimal outcome model, the variation in modeling the propensity score does not affect treatment effect estimates—this suggests that double-robustness holds and the outcome model is correctly specified. The same approach can be taken by holding the propensity score model fixed and varying (and deliberately misspecifying) the outcome model. Wallace et al. (3) demonstrated the ability of this approach to detect deficiencies in model specification in settings where the NUC assumption held, using bootstrapping to gauge the stability of the estimates within the context of their variability. While the authors did not explicitly refer to testing NUC, they did note, “A key property of this approach is that rather than merely an avenue towards model selection, double robustness may be used for model validation” (3, p. 863). This hints at the possibility of using this approach to test for NUC: If estimates remain stable when the analyst varies one model and holds the other fixed, is this evidence that NUC holds? Unfortunately, no: As demonstrated in Web Appendix 3 (data generated accordinging to Web Figure 2; results shown in Web Figures 3 and 4), when there exists an unmeasured confounder, estimates may be stable—but biased. Thus, in any real data analysis, stability alone in a doubly robust estimator is not enough to suggest unbiasedness. Fortunately, not all hope is lost. There exist a number of proposals for sensitivity analyses to determine the potential impact of NUC violations and even to correct for the presence of an unmeasured confounder. It may not be possible to test for or detect NUC, but there are many tools with which to examine the sensitivity of analyses to the violation of this assumption. Monte Carlo approaches focus on repeatedly generating the unmeasured confounder itself; the generated values can then be used directly in an analysis including the “imputed” |$U$| as though it were measured, or generating a bias-correction term that depends on draws from the proposed distribution of |$U$| (4–6). In both cases, the conditional distribution of |$U\!\!\!\! \mid \!\!\! X,A,Y$| must be specified. Alternative sensitivity analyses require the analyst to specify the degree of unmeasured confounding by postulating a value for the coefficient of |$H({\mathrm{\psi}}^{\dagger})$|, |${\mathrm{\alpha}}_1$|, in the propensity score model and then perform estimation with this fixed offset in the propensity score model (7). Yet another approach draws on a pragmatic, decision framework to determine what would be required to reverse an observed association (8, 9). While specifying a conditional distribution for |$U$| or determining what constitutes a plausible degree of confounding bias is not a trivial task, with the advent of large data resources such as electronic health records, estimating these quantities in auxiliary data sources is increasingly feasible. This brings us to our second reason for optimism. It may reassure readers to recall that the impact of NUC violations is greatest when the unmeasured confounder and the measured confounders are uncorrelated—a setting that is arguably unlikely. Consider the reasonable scenario in which the measured and unmeasured confounders are correlated. As shown in Web Appendix 4 (Web Figure 5), the bias associated with omitting a confounder is dramatically smaller when its correlation with the measured confounder is substantial. That is, even when the strength of the unmeasured confounding is high, the inclusion of a measured confounder that is strongly correlated with that unmeasured variable can diminish the bias in estimating |$\mathrm{\psi}$|. While existing hints in the causal inference literature suggest that it might be possible to unmask unmeasured confounders, our simulations remind readers that we can neither formally test for violations of NUC nor use double-robustness to suggest its absence or truly “validate” the models under consideration. These results should reinvigorate us as researchers to engage more fully in scientific collaboration, to ensure that methodologists and data analysts are present when data collection and study design are discussed. Further, the expansion of methods for assessing sensitivity to unmeasured confounding deserves greater attention in a wider array of realistic settings. In spite of these somewhat disappointing results in terms of assessing violations of NUC, there is reassurance to be had: If unmeasured confounders are known or likely to be correlated with those confounders available to the data analyst, this can significantly mitigate the impact of the bias due to violation of the NUC assumption. Conflict of interest: none declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.089 | 0.248 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.003 | 0.021 |
| Scholarly communication | 0.007 | 0.016 |
| Open science | 0.004 | 0.005 |
| Research integrity | 0.010 | 0.015 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".