The epidemiology of two things considered together. Commentary on:<i>Explanation in Causal Inference: Developments in Mediation and Interaction</i>, by Tyler J. VanderWeele
Bibliographic record
Abstract
For much of its history, epidemiology has focused on study designs and statistical analyses for estimating the causal relation between an exposure of interest, X, and a health outcome Y. Other variables were a nuisance in this relationship, acting as confounders or modifiers, but the intent was to identify the consequences of manipulating in isolation just the one exposure X. A conceptual revolution occurred in 1986, however, with the publication of a baroquely methodological tome by Jamie Robins, in which this problem was extended across time to various measures X1, X2, X3, …, Xn.1 Certainly there had been longitudinal data analysis before 1986, but not with an explicitly causal foundation nor a formal approach to handling time-dependent confounding in which a covariate, Z, could be affected by X1, and subsequently be a cause of, and therefore potentially a confounder of, X2.2 This methodological development had wide-ranging implications. One was that it made epidemiology problems inherently structural, because the analysis depended on a specific arrangement of exposures and covariates arranged over time. Robins represented this structure with innovative tree-graphs called FFRCISTGs, but these were never widely adopted by the field. Later, Pearl’s directed acyclic graphs (DAGs) gained wide popularity as a means to convey this structural information.3 Another consequence of Robins’ innovation was the recognition that the new methods would apply equally well if X1 and X2 were the same quantity measured at different times, or if they were two entirely distinct variables measured at different times. This therefore provided a formal approach to mediation and to causal definitions of direct and indirect effects.4 Mediation had a strange history within epidemiology, because for decades it was a method that was ubiquitous in practice but without any formal development or justification in our major textbooks. What was commonly done was to fit a regression model and to take the partial coefficient on the exposure, conditional on baseline confounders, to be an estimate of the total effect of the exposure. Then, a subsequent regression was fit with the addition of a variable presumed to be on the pathway from exposure to outcome, and the new partial coefficient for the exposure in this elaborated model was taken to be an estimate of the (controlled) direct effect, which is to say that part of the total effect that was not relayed through the modelled intermediate. We now know this method as the ‘Baron and Kenny’ approach after a highly influential review article by two social psychologists,5 although it had been a routine method in practice long before this review. In fact, there is even some brief mention of this strategy in a classic epidemiology textbook by Susser in 1973, where he suggested that control for ‘intervening’ variables reduces the conditional effect measure to the null and that this represents a valid strategy for elucidating causal structure in epidemiological research (p. 122).6 The next mention of this strategy in an epidemiology textbook appears to be the first edition of Szklo and Nieto in 2000, which likewise suggested without formal justification or citation that to treat an intermediate variable as a ‘confounder’ (in the sense that statistical control is warranted) is appropriate for mechanistic inference, and that the presence of a residual effect of X on Y would indicate that in addition to affecting the intermediate, the exposure may have a ‘direct toxic effect’ (p.184).7 Robins and Greenland had formally defined direct and indirect effects in terms of potential outcomes in their 1992 paper, but this work seems to have had no discernible impact on epidemiological practice for at least a decade. The numerical example provided in that paper showed that the Baron and Kenny method could be horribly misleading, but there was little general appreciation in the field for the intuition behind this fallacy and how it might be avoided. What finally allowed the message to break through was the advent of the DAG, first presented by Pearl in the statistical literature in 19958 and then brought by Greenland to the annual meeting of the Society for Epidemiologic Research in the summer of 1997. By 1999, Greenland, Pearl and Robins had collaborated on a DAG tutorial for epidemiologists,9 and 3 years later Hernán and Cole provided a DAG-based explanation for the cautionary example in the 1992 paper.10 With wider understanding that the Baron and Kenny approach failed in this instance because of unmeasured confounding between the intermediate and outcome variables, mediation once again started to feel safe for general use. But there was another problem, and this had to do with interaction. The Baron and Kenny method assumes homogeneity of the intermediate-stratum direct effects. When there is interaction between exposure and intermediate in the scale of the regression model, this method fails, as does any additive partition of the total effect into direct and indirect components.11 This impasse was finally sidestepped with the innovation of ‘natural’ effects in which the mediator is fixed not to a single predetermined value, but instead to a potentially unknown value that would be observed under a baseline exposure for each unit.12 The decade from 2005 to 2015 saw a remarkable explosion in work on mediation, including the careful cataloguing of alternative identification assumptions and the extension to a variety of model forms including hazards regression. A large proportion of this tsunami of new insights and approaches came from Tyler VanderWeele, who meticulously worked through every permutation and extension in one paper after another, and then finally structured this body of work into his remarkably comprehensive 2015 textbook.13 The problem of decomposing total effects in the presence of interaction reveals an intrinsic connection between these topics. This fundamental overlap explains VanderWeele’s overall structure for the book, which is to explain mediation, then to explain interaction and then to explain how they reconcile into a single unified framework as the theory of two causes considered together. In an earlier burst of activity around interaction in the 1980s, epidemiologists had sorted out the effects of X and Z when each was potentially a cause of Y but one did not affect the other.14 When the two variables are binary and unconfounded for example, this results in 16 potential outcomes types. But the addition of an arrow from X to Z explodes this number to 64. Robins and Greenland had reduced the complexity of this problem considerably in 1992 by imposing an assumption of monotonicity, which brings the number of types back down to a tractable 18. VanderWeele, however, provides an elegant and complete four-way decomposition of this structure into a proportion due to mediation, a proportion due to interaction, a proportion due to both and a proportion due to neither. It is difficult now to imagine that there is anything remaining still to be said about this topic. Effect decomposition in sociology, via structural equations modelling (SEM), had been a target of much deserved criticism over the years.15 There are two important themes in the VanderWeele book that I believe help to guard against the statistically licentious abuses that embarrassed the SEM literature. The first is a nagging preoccupation with identification. Causal inference one thing at a time is already hard enough, and considering two things at a time is often simply insurmountable. VanderWeele enumerates in each instance the specific assumptions necessary to warrant causal language. For natural direct and indirect effects, these conditions can be especially discouraging. The second theme follows logically from the pessimism surrounding the first: ubiquitous sensitivity analyses. Since we must depend on various independence assumptions to imbue a direct or indirect effect estimate with a causal interpretation, it is important to know how fragile our qualitative interpretations are to realistic violations of these assumptions. VanderWeele has assembled a suite of sensitivity analysis techniques to accompany mediation analyses and he promotes these vigorously. We can never assure that all analytical work will be done honestly and responsibly, but by providing these tools, VanderWeele has at least allowed that any researcher who does want to strive for robust and accountable inference will have the methods at hand to do so. With straightforward prose and many technical details relegated to appendices, this book is clearly designed to be a practical manual for applied researchers. This hands-on focus is supported by frequent examples of software code in SAS and Stata and by worked numerical examples. There is an unavoidable tension that arises when methodologists aim to be accessible to a broad audience of users, however. At first a method is known only to a small group of statistical specialists and substantive applications are rare. Then someone writes a tutorial with code or creates macros for commercial software packages, and suddenly the method takes off, appearing with increasing frequency in papers written by subject matter experts in clinical and public health journals. But these applications are inevitably uneven in quality. As the user base expands, many apply the method without a deep understanding of what they are doing and they can do so because the software packages automate the process for them. This seems an adverse consequence of popularizing new methods, but there is really no alternative. We can’t expect statistical modelling to be restricted to a small cadre of highly trained monks sequestered in an ivory tower. Rather, I think we have to strive to do exactly like VanderWeele: present the method as transparently and honestly as possible, show people the assumptions that need to be verified and show them how to make these checks, and then give them sensitivity analysis tools to investigate violations of the assumptions. The seduction of ‘powerful’ methods will inevitably lead some users to excessively wishful thinking in order to surmount challenges to identification and estimation. There is nothing to be done about this other than to critique these studies when they are published, and slowly educate the field toward better practice over time. In the 1970s, paleontologists proposed a theory of ‘punctuated equilibria’ which suggested that species were static for long stretches of evolutionary time, until every so often rapid changes burst forth. Innovation in epidemiological methods seems to follow a similar ebb and flow. After a flurry of activity in the 1980s on interaction analyses to identify causal mechanisms, a pessimistic paper in 1991 pretty much ended the discussion.16 No significant new innovation appeared for about 15 years, until the topic was taken up by VanderWeele and finally interwoven intimately with the more recent developments in mediation. In hindsight, this fundamental connection seems obvious but it awaited the methodical development of a counterfactual model for mediation to see the relation with earlier theories of synergism based in sufficient component causes and latent causal strata. VanderWeele has excelled in making such connections, first laying out these two topics side by side with encyclopaedic breadth and then neatly fusing them into an elegant synthesis. This work now defines the comprehensive toolkit for the epidemiology of two causally ordered factors considered jointly. It remains for us to incorporate these developments into our teaching, for authors to apply these methods in practice and for reviewers and editors to assess the quality of this work and to enforce VanderWeele’s emphasis on justifying identification conditions, and where these are in doubt, reporting appropriate sensitivity analyses. Funding was received from the Canada Research Chairs program. Conflict of interest: None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.121 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.006 | 0.011 |
| Scholarly communication | 0.006 | 0.014 |
| Open science | 0.007 | 0.004 |
| Research integrity | 0.080 | 0.102 |
| Insufficient payload (model declined to judge) | 0.008 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".