Causal Inference and Confounding: A Primer for Interpreting and Conducting Infectious Disease Research
Bibliographic record
Abstract
Causal inference has become a mainstream branch of statistics [1], changing the landscape of analysis and interpretation in many fields, including infectious disease research. A common research objective is to infer (or test for) the impact of an intervention on an outcome using study data. Sometimes this is accomplished through a randomized trial, in which case the conclusion of the trial can often be interpreted causally, as treatment allocation is independent of any factors that predict the outcome. Randomization ensures that the (2 or more) treatment groups are “fair” and do not differ systematically in any way that might lead to one randomization group having better outcomes than the other at the outset of the study. In this randomized setting, any difference in outcome between the groups can be attributed to the treatment, because that is the only way in which the randomization groups are systematically different. Note that this is an idealized randomized trial. The benefits of randomization can be compromised if participants in a trial do not adhere to their assigned treatment, or systematically behave differently after being assigned to their randomization group (eg, if individuals in the arm assigned to a placebo in a trial of an analgesic systematically took more over-the-counter medications than those in the active drug arm). There are many settings where randomization is not possible, yet we would still like to understand if a particular factor causes changes in the outcome. Perhaps the factor of interest is thought to be harmful, so there are ethical concerns around assigning participants to receive it. Alternatively, the population of interest may be small, and so it is not feasible to conduct a randomized trial each time a potential factor or treatment of interest is identified. Consider, for example, a study of 2 different Food and Drug Administration-approved immunosuppressants in patients undergoing allogeneic hematopoietic cell transplantation. The treatment itself could be subject to randomization, but the population is small enough that hospitals or ethics board may be unwilling to support a trial for already-approved drugs in this group. As an alternative, a researcher could conduct an analysis of an observational, or nonexperimental, dataset. In these data, treatments are not assigned by randomization but rather by the choice of individual treating physicians. In this case, one treatment, perhaps longer in therapeutic use, may be given to patients with lower risk of poor outcomes whereas patients with higher risk (eg, test positive for cytomegalovirus infection, or poorer donor-recipient matching) receive the newer, more aggressive therapy. If different outcomes are observed between these 2 treatment groups, can we conclude that it is because of the treatment? Or might the difference be attributable to the pretreatment differences in risk between the groups? This situation, where there is a common cause for both the treatment allocation and the outcome, is called confounding. Confounding can distort or bias the observed treatment effect, and is common in observational studies, where patients with a worse prognosis preferentially receive or choose one intervention over another. Other forms of bias may also be present. Selection bias (also called participation bias) occurs when those individuals included in a study may differ systematically from the population of interest. This could arise if the analytic dataset is drawn from a particular health insurer: doing so could exclude older individuals and individuals of lower incomes or who are unemployed or unhoused. Causal inference is a discipline that attempts to ascribe a causal relationship between an intervention or exposure of interest and an outcome based on data. The discipline encompasses many different estimation approaches, unified by a general approach that emphasizes precision at each stage of research, from the design to the method of analysis. The design side focuses on clear definitions of the treatment, the outcome, and inclusion/exclusion criteria, whereas the analysis side aims to correct for biases due to imbalances (lack of “fairness” between comparison groups) that can distort the true impact of a treatment. Causal inference does not refer to any single method of analysis; rather, the analysis step can encompass a range of different methods and tools. The validity of the causal interpretations depends on the plausibility of the assumptions that are made regarding the existing of confounding and the faithfulness of the analysis model to the underlying data-generating process (“nature”). As outlined by Goetghebeur et al [2], the explicit formalization of all definitions, the target of the causal effect, and the method of estimation with all the associated assumptions on data availability are critical to being able to assert a true causal relationship (or indeed lack thereof) and not merely a perhaps-biased association. See Box 1 for a simplified causal analysis workflow. A central tenet in causal analysis is to focus on a single treatment or exposure [3]. Causal graphs, also called directed acyclic graphs, are useful tools for encoding beliefs on proposed relationships between variables and to center the analyses on a single treatment or exposure (Figure 1). Example of a causal diagram (also called a directed acyclic graph) showing a simple setting with a treatment (eg, standard vs nonstandard immunosuppressant) whose effect on the outcome is confounded. In these diagrams, the direction of an arrow is indicative of a causal relationship, where changing the level or value of the variable at the origin of the arrow will result in changes in the variable to which it points. A, Cytomegalovirus (CMV) status is a confounding variable (ie, a common cause of both treatment and the outcome) for the treatment-exposure relationship. Here, if CMV is the only confounder, then adjusting for CMV status is sufficient to remove confounding bias so that a causal interpretation can be drawn. In generic notation, we can depict this scenario as in (B). In general, there may be multiple confounders that must be taken into account. Note that for each potential treatment or factor of interest, a new causal diagram may be needed to ensure that the analysis captures the confounders (common causes) relevant to a particular treatment-outcome pair. Key steps for a causal analysis. Simplified and adapted from Goetghebeur et al [1]. Nguyen et al’s [4] comparison of the effect of tenofovir disoproxil (TDF) treatment for chronic hepatitis B virus infection on the incidence of hepatocellular carcinoma provides an example of causal principles. The authors clearly defined the treatment comparison (newer medication vs none) and the target population (Asian population with chronic hepatitis B), and addressed confounding via matching, an approach that selects an analytic sample of individuals in each treatment group who are similar in terms of any potential confounders. The initial pool of data available to the authors showed that treated individuals were much more likely to have liver cirrhosis at baseline and have higher liver enzyme activity (ie, evidence of confounding). Following the construction of the matched analytic dataset, an analytic subsample was created in which 2 treatment groups were comparable with respect to these and other important covariates. The comparison between the treatment groups (new treatment vs none) was thus rendered fairer, and the authors concluded that TDF treatment was associated with significantly reduced incidence of hepatocellular carcinoma in this cohort. While causal inference provides a framework and guidance for ascribing causal interpretation to an estimated association, it is not suited to all research questions. For instance, causal inference is not used for forecasting (eg, prediction of hospital admissions for respiratory viruses) or building a clinical decision-support tool to predict probable diagnosis based on a collection of symptoms. Nevertheless, understanding the aim of fairness in comparisons that underlies causal inference is critical to infectious disease researchers wishing to ascribe more than an associational interpretation to their work, and to better judge the validity of causal claims in the literature. Financial support. This work was supported by a Discovery Grant from the Natural Sciences and Engineering Research Council of Canada [NSERC RGPIN-2019-04230]. E. E. M. M. is supported by a Chercheur de Mérite Career Award from the Fonds de Recherche du Québec, Santé [FRQ-S 309780] and holds a Canada Research Chair in Statistical Methods for Precision Medicine from the Canadian Institutes of Health Research [950-233182 X-257241].
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.158 | 0.280 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.007 |
| Bibliometrics | 0.009 | 0.007 |
| Science and technology studies | 0.003 | 0.023 |
| Scholarly communication | 0.012 | 0.011 |
| Open science | 0.008 | 0.007 |
| Research integrity | 0.009 | 0.020 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".