MétaCan
Menu
Back to cohort
Record W4220703773 · doi:10.1093/ije/dyac056

The quintessence of causal DAGs for immortal time bias: time-dependent models

2022· article· en· W4220703773 on OpenAlexaff
Ian Shrier, Samy Suissa

Bibliographic record

VenueInternational Journal of Epidemiology · 2022
Typearticle
Languageen
FieldMathematics
TopicAdvanced Causal Inference Techniques
Canadian institutionsMcGill UniversityJewish General Hospital
Fundersnot available
KeywordsQuintessenceCausal inferenceCausal modelCausality (physics)MedicinePsychologyEconometricsStatisticsMathematicsPhysicsCosmology

Abstract

fetched live from OpenAlex

We believe the recent editorial1 on causal directed acyclic graphs (causal DAGs) for immortal time bias (ITB) is inaccurate for the example of the effect of heart transplant versus medical treatment on mortality (Figure 1, adapted from their Figure 2). Figure 1a–d suggests ITB occurs because there is either an unmeasured common cause (U) of (i) the measured exposure (A*) and the outcome (Y) (measurement error), or (ii) a variable causing both exclusion of participants (E = 0) and the outcome (collider stratification bias). However, these DAGs do not represent the essential time-dependent nature of ITB. (adapted from Mansournia et al.1). Causal diagrams according to the previous editorial which represent different approaches for handling immortal times. U is an unmeasured variable, A is the true value exposure, A* [in (a)] is the value of exposure used in the analysis, Y is the outcome, and E [in (b)] is exclusion of immortal time records. Figure 1d represents a sequentially randomized trial where Am is the month of transplant, and C = 0 indicates participant-time after Am is excluded if a transplant was not received by Am. These causal diagrams suggest immortal time bias is due to unmeasured confounding in (a), collider stratification bias due to conditioning on E = 0 in (b), and no bias in a time-dependent analysis in (c) or (d) if A does not cause Y. If in reality A did cause Y (i.e. if there was an additional arrow from A to Y, then the modified (c) and (d) would suggest there would be confounding bias even with time-dependent analyses. Panel (a) uses recommended notation (see text) to illustrate the causal diagram for immortal time bias in a randomized study where the true exposure (A) may vary over time and be different early in the study (A0) and at subsequent time points. We only show one additional time point for simplicity (A1). The outcome is measured after baseline but before time point 1 (Y0+), and after time point 1 (Y1+). We include A* from the previous editorial as a composite measure of exposure that some investigators have used in analyses. In the context of the heart transplant study described in the editorial, treatment early in the study (A0) is a cause of survival/death before time point 1 (Y0+), and if there are delayed effects, it is also a cause of survival/death after time point 1 (Y1+). Treatment later in the study (A1) is a cause of survival/death at Y1+. Although not usually included in causal directed acyclic graphs (DAGs) because they are not common causes of exposure and outcome, we include other causes of death at Y0+ (U0) and Y1+ (U1) because they help illustrate the mechanism of immortal time bias. The arrow from A0 to A1 indicates that a participant who had their transplant early in the study before time point 1 also had their transplant after time point 1. The arrow from treatment assignment to A1 illustrates that participants assigned surgery, who did not receive surgery early in the study, will receive surgery at a subsequent time point if they are alive. The arrow (dotted for emphasis, and because it is partially deterministic) from Y0+ to A1 represents the fundamental reason for ITB; in order to have a transplant at time point 1, one must have Y0+ (outcome early in study) = 0. There is no arrow from A* to Y0+ or Y1+ because the composite variable for exposure does not ‘cause’ anything. Because the analysis conditions on Y0+, a non-causal association is created between A0 and U0, leading to a biased estimate for the effect of A on Y. Panel (b) represents the corrected causal DAG for the ‘selection bias’ ITB example from the editorial, where participants are included if they (i) only received medical treatment (A0 = A1 = 0), or (ii) lived to have surgery (Y0+ = 0 and A1 = 1). To produce accurate causal statements involving time-varying quantities like person-time, nodes in causal DAGs must represent random variables at a particular point in time.2,3 Thus, total person-time at risk should not be used as a node. Further, the editorial’s ‘causal DAGs’ are inconsistent with the described data-generating process:3,4 The value of A* depends on knowing that participants are alive at any time after cohort entry (‘higher probability of surviving the waiting period’). Therefore, A* is a function of (caused by) Y at the end of the waiting period. This essential component is missing from the causal DAG. Instead, U is included as a common cause of A* and Y. These issues mean the proposed causal DAG misses the actual causal mechanism underlying ITB. Figure 2a illustrates a more accurate causal DAG for the ITB-induced misclassification error example (Figure 1a). We use recommended notation where both exposure and outcome are correctly specified as time-dependent variables (Vt: variable measured at time t).2,3,5 For illustrative purpose, we use two time points, 0 and 1, with follow-up continuing beyond time 1. Other causes of the outcome between baseline and time point1 (Y0+) and after time point1 (Y1+) are indicated by U0 and U1, respectively. The DAG also includes A*, the incorrectly measured time-independent exposure, to be consistent with the editorial. The dotted arrow from Y0+ to A1 is partially deterministic and represents the fact that participants who survive the waiting period can be exposed at A1 but those who die cannot. This critical difference illustrates why measuring and conditioning on the common cause U of A* and Y in Figure 1a will not remove ITB. Similar limitations apply to the editorial’s ‘selection bias’ example (Figure 1b), where ‘immortal exposure time’ from participants in the transplant group is excluded if they receive a transplant. Figure 2a also illustrates that the foundational cause of immortal time bias in both exposure misclassification and selection bias is collider stratification bias, not measurement error. In Figure 2a, conditioning on Y0+ opens a non-causal path (A0-Y0+-U0), regardless of whether there is a confounder of the A-Y relationship that is conditioned on, and whether variables are considered time-dependent. The only required modification for the editorial’s selection bias example is the addition of an arrow from Y0+ to a selection node (S) (Figure 2b). This indicates that we only include participant person-time from those (i) assigned exclusively medical treatment (Treatment assigned = 0), or ii) assigned surgical treatment (Treatment assigned = 1) if they survive to receive transplant (Y0+= 0), thus excluding their prior medical treatment time. The different inclusion criteria for participants assigned medical versus surgical treatment means the critical exchangeability assumption is likely violated. Therefore, to avoid immortal-time biases, the study design or analysis needs to consider explicitly that exposure is time-varying, and use an appropriate analysis.6–8 Under the special context when A does not affect Y (Figure 1c, d), Y is not a collider and there is no ITB with time-dependent analyses. Using a composite non-time-dependent measure of exposure like A* will likely remain biased even if A does not affect Y. The editorial says that most ITB examples described by Suissa9,10 are due to similar mechanisms. Although possible, there are some important nuances that likely lead to different causal diagrams. Suissa’s many examples of ITB include Y0+ causing A1 (i) only in the medical group (time-based and event-exposure based cohort9), (ii) only in the treatment group (event-based and multiple event-based cohort9), or (iii) both (time-based and event-based cohorts9). Exposure-based cohorts9 are slightly different because Y0+ affects A1 through its effects on the value of start time used, T0. No such approval is required for a letter to editor that does not describe new original research findings. The authors would like to thank Robert Platt, Jamie Robins and Sander Greenland for their very helpful comments on early drafts. We would especially like to thank Sander Greenland for suggesting important corrections to our own earlier causal DAGs for immortal time bias. Both authors contributed to the writing of this manuscript. None declared.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.011
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.835
Threshold uncertainty score0.997

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0060.011
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.281
GPT teacher head0.470
Teacher spread0.189 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of EpidemiologySame topicAdvanced Causal Inference TechniquesFrench-language works237,207