Genetic instruments with too many strings: acknowledging pleiotropy and population structure in Mendelian randomization studies
Notice bibliographique
Résumé
This commentary refers to ‘Genetically modulated educational attainment and coronary disease risk’, by L. Zeng et al., pages 2413–2420. The availability of large genetic databases has prompted renewed interest in genetic contributions to educational attainment and the effects of education on health. In a recent issue of the European Heart Journal, Zeng et al.1 use cohorts from across Europe and the USA to investigate education-associated genotypes and their role in coronary artery disease (CAD). We would like to commend the authors for raising important questions about the relationships between genetics, education, and CAD. Nonetheless, it is our belief that the methods employed have important limitations not discussed by the authors. Here, we note three of these limitations. First, the authors investigate whether the genetic risk score (GRS) for education was associated with CAD independently of its effect on education by adjusting for number of years of attained education in the regression of CAD on the GRS. The logic is to ‘block’ the effect of the GRS through attained education (Figure 1). However, beyond ‘blocking’ the effect, this method will also create a new correlation between GRS and all other causes of education because it becomes a ‘collider’2 (Figure 1, dashed arrow). Therefore, even if the effect of the education GRS on CAD is only mediated by years of education, the GRS will continue to be correlated with CAD after adjustment for years of education. Any residual association after adjusting for years of education cannot then be interpreted as an independent or pleiotropic effect of the GRS. To answer this question by adjustment, appropriate mediation methods would be required.3 Adjusting for years of education blocks the effect of the genetic risk score on coronary artery disease but creates a correlation between the genetic risk score and the confounders of education and coronary artery disease (dashed line). Second, one of the three major assumptions of Mendelian randomization (MR) analyses is that the effect of the single-nucleotide polymorphisms on the outcome (CAD) only passes through the exposure (education).4 Therefore, if the authors’ claim to have found pleiotropic effects of the education GRS on CAD were, in fact, true, this would also imply that the results of the MR analysis are biased. Therefore, it is difficult to know what to conclude. If the education GRS is pleiotropic, as the authors claim, then the main MR analysis must be biased. Alternatively, if the authors claim that the MR analysis is valid, this implies that there is no pleiotropic effect of the education GRS on CAD. Lastly, confounding by population stratification is difficult to control when using a GRS, particularly, with geographically patterned phenotypes, as in the case of education and heart disease.5 The authors use five principal components to adjust for this but it has been shown that geographic confounding may not be successfully controlled even with 40 principal components. This confounding would invalidate both the primary and MR analyses. Furthermore, such residual confounding is consistent with the finding that GRS-CAD associations become null when adjusting for body mass index and smoking, i.e. adjusting for these factors help account for population stratification thus showing the independent association with GRS to be null. Conflict of interest: none declared.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,031 | 0,197 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,003 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,004 | 0,010 |
| Communication savante | 0,005 | 0,008 |
| Science ouverte | 0,004 | 0,003 |
| Intégrité de la recherche | 0,065 | 0,067 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».