MétaCan
Menu
Retour à la cohorte
Enregistrement W4375951527 · doi:10.1093/infdis/jiad144

Causal Inference and Confounding: A Primer for Interpreting and Conducting Infectious Disease Research

2023· article· en· W4375951527 sur OpenAlexafffund
Erica E. M. Moodie

Notice bibliographique

RevueThe Journal of Infectious Diseases · 2023
Typearticle
Langueen
DomaineMathematics
ThématiqueAdvanced Causal Inference Techniques
Établissements canadiensMcGill University
Organismes subventionnairesNatural Sciences and Engineering Research Council of CanadaCanadian Institutes of Health Research
Mots-clésCausal inferenceConfoundingInferenceInfectious disease (medical specialty)Primer (cosmetics)MedicineDiseaseBiologyComputational biologyVirologyComputer scienceInternal medicineArtificial intelligencePathologyChemistry

Résumé

récupéré en direct d'OpenAlex

Causal inference has become a mainstream branch of statistics [1], changing the landscape of analysis and interpretation in many fields, including infectious disease research. A common research objective is to infer (or test for) the impact of an intervention on an outcome using study data. Sometimes this is accomplished through a randomized trial, in which case the conclusion of the trial can often be interpreted causally, as treatment allocation is independent of any factors that predict the outcome. Randomization ensures that the (2 or more) treatment groups are “fair” and do not differ systematically in any way that might lead to one randomization group having better outcomes than the other at the outset of the study. In this randomized setting, any difference in outcome between the groups can be attributed to the treatment, because that is the only way in which the randomization groups are systematically different. Note that this is an idealized randomized trial. The benefits of randomization can be compromised if participants in a trial do not adhere to their assigned treatment, or systematically behave differently after being assigned to their randomization group (eg, if individuals in the arm assigned to a placebo in a trial of an analgesic systematically took more over-the-counter medications than those in the active drug arm). There are many settings where randomization is not possible, yet we would still like to understand if a particular factor causes changes in the outcome. Perhaps the factor of interest is thought to be harmful, so there are ethical concerns around assigning participants to receive it. Alternatively, the population of interest may be small, and so it is not feasible to conduct a randomized trial each time a potential factor or treatment of interest is identified. Consider, for example, a study of 2 different Food and Drug Administration-approved immunosuppressants in patients undergoing allogeneic hematopoietic cell transplantation. The treatment itself could be subject to randomization, but the population is small enough that hospitals or ethics board may be unwilling to support a trial for already-approved drugs in this group. As an alternative, a researcher could conduct an analysis of an observational, or nonexperimental, dataset. In these data, treatments are not assigned by randomization but rather by the choice of individual treating physicians. In this case, one treatment, perhaps longer in therapeutic use, may be given to patients with lower risk of poor outcomes whereas patients with higher risk (eg, test positive for cytomegalovirus infection, or poorer donor-recipient matching) receive the newer, more aggressive therapy. If different outcomes are observed between these 2 treatment groups, can we conclude that it is because of the treatment? Or might the difference be attributable to the pretreatment differences in risk between the groups? This situation, where there is a common cause for both the treatment allocation and the outcome, is called confounding. Confounding can distort or bias the observed treatment effect, and is common in observational studies, where patients with a worse prognosis preferentially receive or choose one intervention over another. Other forms of bias may also be present. Selection bias (also called participation bias) occurs when those individuals included in a study may differ systematically from the population of interest. This could arise if the analytic dataset is drawn from a particular health insurer: doing so could exclude older individuals and individuals of lower incomes or who are unemployed or unhoused. Causal inference is a discipline that attempts to ascribe a causal relationship between an intervention or exposure of interest and an outcome based on data. The discipline encompasses many different estimation approaches, unified by a general approach that emphasizes precision at each stage of research, from the design to the method of analysis. The design side focuses on clear definitions of the treatment, the outcome, and inclusion/exclusion criteria, whereas the analysis side aims to correct for biases due to imbalances (lack of “fairness” between comparison groups) that can distort the true impact of a treatment. Causal inference does not refer to any single method of analysis; rather, the analysis step can encompass a range of different methods and tools. The validity of the causal interpretations depends on the plausibility of the assumptions that are made regarding the existing of confounding and the faithfulness of the analysis model to the underlying data-generating process (“nature”). As outlined by Goetghebeur et al [2], the explicit formalization of all definitions, the target of the causal effect, and the method of estimation with all the associated assumptions on data availability are critical to being able to assert a true causal relationship (or indeed lack thereof) and not merely a perhaps-biased association. See Box 1 for a simplified causal analysis workflow. A central tenet in causal analysis is to focus on a single treatment or exposure [3]. Causal graphs, also called directed acyclic graphs, are useful tools for encoding beliefs on proposed relationships between variables and to center the analyses on a single treatment or exposure (Figure 1). Example of a causal diagram (also called a directed acyclic graph) showing a simple setting with a treatment (eg, standard vs nonstandard immunosuppressant) whose effect on the outcome is confounded. In these diagrams, the direction of an arrow is indicative of a causal relationship, where changing the level or value of the variable at the origin of the arrow will result in changes in the variable to which it points. A, Cytomegalovirus (CMV) status is a confounding variable (ie, a common cause of both treatment and the outcome) for the treatment-exposure relationship. Here, if CMV is the only confounder, then adjusting for CMV status is sufficient to remove confounding bias so that a causal interpretation can be drawn. In generic notation, we can depict this scenario as in (B). In general, there may be multiple confounders that must be taken into account. Note that for each potential treatment or factor of interest, a new causal diagram may be needed to ensure that the analysis captures the confounders (common causes) relevant to a particular treatment-outcome pair. Key steps for a causal analysis. Simplified and adapted from Goetghebeur et al [1]. Nguyen et al’s [4] comparison of the effect of tenofovir disoproxil (TDF) treatment for chronic hepatitis B virus infection on the incidence of hepatocellular carcinoma provides an example of causal principles. The authors clearly defined the treatment comparison (newer medication vs none) and the target population (Asian population with chronic hepatitis B), and addressed confounding via matching, an approach that selects an analytic sample of individuals in each treatment group who are similar in terms of any potential confounders. The initial pool of data available to the authors showed that treated individuals were much more likely to have liver cirrhosis at baseline and have higher liver enzyme activity (ie, evidence of confounding). Following the construction of the matched analytic dataset, an analytic subsample was created in which 2 treatment groups were comparable with respect to these and other important covariates. The comparison between the treatment groups (new treatment vs none) was thus rendered fairer, and the authors concluded that TDF treatment was associated with significantly reduced incidence of hepatocellular carcinoma in this cohort. While causal inference provides a framework and guidance for ascribing causal interpretation to an estimated association, it is not suited to all research questions. For instance, causal inference is not used for forecasting (eg, prediction of hospital admissions for respiratory viruses) or building a clinical decision-support tool to predict probable diagnosis based on a collection of symptoms. Nevertheless, understanding the aim of fairness in comparisons that underlies causal inference is critical to infectious disease researchers wishing to ascribe more than an associational interpretation to their work, and to better judge the validity of causal claims in the literature. Financial support. This work was supported by a Discovery Grant from the Natural Sciences and Engineering Research Council of Canada [NSERC RGPIN-2019-04230]. E. E. M. M. is supported by a Chercheur de Mérite Career Award from the Fonds de Recherche du Québec, Santé [FRQ-S 309780] and holds a Canada Research Chair in Statistical Methods for Precision Medicine from the Canadian Institutes of Health Research [950-233182 X-257241].

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,158
score de la tête « metaresearch » (Gemma)0,280
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Méthodes · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,842
Score d'incertitude au seuil0,837

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,1580,280
Méta-épidémiologie (sens strict)0,0030,003
Méta-épidémiologie (sens large)0,0060,007
Bibliométrie0,0090,007
Études des sciences et des technologies0,0030,023
Communication savante0,0120,011
Science ouverte0,0080,007
Intégrité de la recherche0,0090,020
Charge utile insuffisante (le modèle a refusé de juger)0,0070,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,271
Tête enseignante GPT0,499
Écart entre enseignants0,228 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeThéorique ou conceptuel
DomaineMéthodes
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations7
Publié2023
Routes d'admission2
Résumé présentnon

Explorer davantage

Même revueThe Journal of Infectious DiseasesMême sujetAdvanced Causal Inference TechniquesTravaux en français237 207