Commentary: Reporting and assessing evidence for interaction: why, when and how?
Bibliographic record
Abstract
Thompson1 noted in 1991 that, although more than a decade had passed since the first discussion of ‘interaction’ in the epidemiological literature, debate had by then subsided and few clear conclusions had emerged. He concluded: Unfortunately, choice among theories of pathogenesis is enhanced hardly at all by the epidemiological assessment of interaction … What few causal systems can be rejected on the basis of observed results would provide decidedly limited etiological insight. One could point to several developments that have contributed to this reawakening of interest. The first of these has arisen from recent work by the statisticians in ‘causal modelling’, which has led to new, deterministic, definitions of mechanistic interaction. One such approach derives from the earlier work of Rothman4 on component causes, whereas the other is based on the currently popular focus on counterfactual outcomes. In the former approach, interaction between two causal factors is defined as the presence of a sufficient cause, which involves both factors as component causes, whereas from the latter viewpoint, interaction implies the existence of people in the population who would not have developed disease unless they had been exposed to both factors. These viewpoints have been shown to be mutually consistent5–7 and, subject to certain assumptions, are also consistent with statistical definition of interaction as deviation from additivity of effects on risk. Broadly, the same conclusions follow from stochastic models for independent causes, originally discussed by Rothman8 and Miettinen9 and recently revisited from the standpoint of directed acyclical graphical models by VanderWeele and Robins.10 In common with many discussions of interaction and mechanism, Boffetta et al.3 and Knol and VanderWeele2 ignored the role of time (particularly age). For chronic degenerative diseases, this must be considered in any analysis, but it is not straightforward. Greenland and Poole5 briefly discussed this problem in the context of stochastic sufficient cause models, pointing out that most commonly used additive models impose rigid constraints on the form of cause-specific hazard functions. Such constraints are avoided in the non-parametric additive hazards model of Aalen11 (extended to case–control data by Borgan and Langholz12). However, the time-dependent aspect of the problem seems to have been ignored in deterministic causal theories. Although recent work has established a more rigorous basis for the idea of mechanistic interaction, its identification in many circumstances with non-additivity for effects on risk had already been recognized for many years before Thompson came to his pessimistic conclusion. We can only assume that he regarded establishment of the existence of such interaction as providing ‘decidedly limited etiological insight’. So, what has changed to re-ignite epidemiologists’ enthusiasm for interaction? The answer must surely be the greater availability of genetic data. Thus, discussion of gene–gene interaction often identified with the earlier (and better defined) concept of epistasis pervades the recent genetic epidemiology literature, and few epidemiological grant applications now fail to identify the establishment of ‘gene–environment interaction’ as a primary aim. Yet much of this discussion is as careless in its use of terms as the early epidemiological literature that first prompted debate about the topic 40 years ago; an interaction trumpeted in the title of a paper more often than not turns out to represent a significant interaction term in a logistic regression model, despite the fact that this, in general, has few implications for mechanism. Statistical interaction between two genes in the logistic regression model is often described as epistasis, although the multiplicative model for joint effects of two genes, which is approximately equivalent to a ‘main effects’ logistic model, has previously been used as model for epistasis.13,14 Against this background, perhaps there is a place for the publication of some guidelines. Perhaps, the most important service that could be provided by the issuing of guidelines would be to foster greater clarity in presentation of precisely ‘why’ the ‘interaction’ reported is of interest. As is argued by Knol and VanderWeele2 this, in turn, has implications for ‘how’ it should be reported. In this respect, Boffetta et al.3 contribute little, after the now familiar agonizing about how difficult it is to define ‘interaction’, finishing with a conclusion remarkably similar to Thompson's of 20 years before: It is not always clear whether or not sensible biological conclusions can come from a particular statistical formula of interaction. It is often argued that, regardless of any direct causal interpretation, interaction on the additive scale is important from a public health point of view. Even this is over-stated; as pointed out by Clayton and McKeigue,16 the argument for targeted intervention does not only depend on non-additivity of effects—the ‘prevention paradox’17 remains under multiplicative models for accrual of risk. Despite these reservations, presence or absence of statistical interaction is an important part of any data analysis as it is related to the goodness of fit of the model adopted for risk (either explicitly or implicitly) and, therefore, to the accurate assessment of risks for different risk factor profiles. This brings us to the question of ‘when’ statistical interactions should be reported. The usual criterion for this is Occam's razor, applied either by a classical significance test for interaction or by some optimal prediction criterion such as the Akaike Information Criterion.18 As has been widely recognized, interaction parameters are estimated precisely much less than the main effects so that they are only rarely significant in a single study. However, Boffetta et al.3 are concerned with systematic review, which requires that studies are reported in such a way that evidence can be combined across studies. This raises the possibility that interaction could be insufficiently convincing to report in an individual study, but statistically significant when evidence is combined over studies. Boffetta et al.3 recommend publication of web supplementary tables of data on gene–environment joint effects and, since this recommendation comes at the end of a discussion of the dangers of selective reporting, we can only assume that they are proposing that Occam's razor should not apply in this context. But, clearly reporting of all non-significant interactions in the expectation of later meta-analysis is not feasible. Although not explicitly suggested, it is implied that data on interactions should be presented regardless of their significance when there is sufficient prior expectation that interaction is present. However, in the absence of clear relationship between plausible mechanisms and statistical interaction in a given risk model, how should such prior expectations be informed? Ultimately, there would seem to be no solution to this problem; meta-analysis with an emphasis on joint effects of two or more factors will, of necessity, require assembly of the relevant raw data. The problem of how to report significant interactions remains, and was discussed in some detail by Knol and VanderWeele.2 This essentially concerns the choice of parametrization of an interaction term in a (generalized) linear model. In the case of two binary risk factors, Knol and VanderWeele distinguish between an ‘interaction’ parametrization, which considers the effect of each risk factor combination with one combination taken as baseline, and an ‘effect modification’ parametrization, which reports the effect of one factor for each level of the other. Both of these parametrizations include one or both of the main effects, and address slightly different questions. It is arguable whether it is necessary to publish guidelines that suggest how authors should lay out their tables to best make their point but, nevertheless, this article contains some useful insights. However, it is worth noting that the ‘interaction’ parametrization results in parameter estimates, which are strongly intercorrelated as a result of the use of a shared baseline. This can be misleading and seriously limits subsequent use of results in meta-analyses. In this context, perhaps the idea of ‘floating absolute risk’19deserves mention. Although this has been criticized,20 its limitations are now better understood and presentation of the ‘quasi-variances’21 will often be helpful. Although reporting and assessment of evidence for interaction are not the same thing, they are clearly linked. Opinions will, no doubt, differ as to whether or not the publication of guidelines to aid either process is necessary or even useful. In addressing reporting, Knol and VanderWeele2 have at least clearly defined what they mean by ‘interaction’ and have not made unrealistic claims for its interpretation. In contrast, the guidelines for assessment of evidence suggested by Boffetta et al.3 are in danger of adding to the continuing confusion. Wellcome Trust Principal Research Fellowship (to D.C.); Juvenile Diabetes Research Foundation (to D.C.). The Cambridge Institute for Medical Research is in receipt of a Wellcome Trust Strategic Award (079895). Conflict of interest: None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.039 | 0.214 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.007 | 0.008 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.008 | 0.003 |
| Research integrity | 0.110 | 0.078 |
| Insufficient payload (model declined to judge) | 0.007 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".