MétaCan
Menu
Retour à la cohorte
Enregistrement W2914988643 · doi:10.1093/biomet/asy061

Discussion of ‘Gene hunting with hidden Markov model knockoffs’

2018· article· en· W2914988643 sur OpenAlexfundaboutno aff
Sean Jewell, Daniela Witten

Notice bibliographique

RevueBiometrika · 2018
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGenetic Associations and Epidemiology
Établissements canadiensnon disponible
Organismes subventionnairesNational Institute of Biomedical Imaging and BioengineeringNational Institute of General Medical SciencesNational Institute on Drug AbuseNatural Sciences and Engineering Research Council of CanadaNational Institutes of HealthNational Science Foundation
Mots-clésHidden Markov modelMathematicsMarkov chainMarkov modelArtificial intelligenceStatisticsComputer science

Résumé

récupéré en direct d'OpenAlex

We wish to congratulate the authors on their extension of model-X knockoffs beyond the Gaussian setting considered in Candès et al. (2018) to the well-studied class of Markov chain and hidden Markov models. This contribution builds upon a beautiful body of work that allows us to rethink false discovery rate control in the context of large and complex datasets (Barber & Candès, 2015, 2018; Weinstein et al., 2017; Barber et al., 2018; Candès et al., 2018; Katsevich & Sabatti, 2018). We believe that this innovation will help model-X knockoffs to be used in practical applications, such as genome-wide association studies explored in the paper, as well as other biological problems where Markov chain or hidden Markov models are common, see Sesia et al. (2019, §1.3) for some examples. In this discussion, we comment on two directions that merit further investigation. Sesia et al. (2019) provide an algorithm for generating model-X knockoffs without resorting to the Gaussian approximation of Candès et al. (2018), in the setting where the covariates are generated from a Markov chain or a hidden Markov model. This raises an intriguing question: if the covariates are distributed according to a Markov chain or hidden Markov model, then what could go wrong if we instead generate knockoffs according to the Gaussian approximation in Candès et al. (2018)? In practice, how bad would it be if we were to generate knockoffs the wrong way? Sesia et al. (2019) begin to address this issue within the context of the Crohn’s disease data in §7 of their paper, by comparing the set of variables selected when knockoffs are generated using a hidden Markov model with the set of variables selected when knockoffs are generated using the Gaussian approximation from Candès et al. (2018). They find that generating knockoffs under the hidden Markov model results in more discoveries than using the Gaussian approximation, and on this basis they claim that knockoffs generated under the hidden Markov model have higher power. However, ground truth data is not available for the Crohn’s disease data: the true positive single nucleotide polymorphisms are unknown. Therefore, we propose to study the false discovery rate and power on synthetic datasets. Covariates were generated under a Markov chain model, and knockoffs were generated using (a) the approach proposed in Sesia et al. (2019) and (b) the Gaussian approximation of Candès et al. (2018). Each panel displays the pairwise correlations of the covariates, |$\mathrm{cor}(X_{j}, X_{j-1})$|⁠, versus the pairwise correlations of the knockoffs, |$\mathrm{cor}(\tilde{X}_{j}, \tilde{X}_{j-1})$|⁠, for |$j=2,\ldots, p$|⁠. The pairwise exchangeability condition (1) appears to be violated by the knockoffs generated using the Gaussian approximation. In practice, however, there is little difference between the knockoffs of Sesia et al. (2019) and those of Candès et al. (2018) in terms of power and false discovery rate control; see Fig. 2. In fact, the Gaussian approximation of Candès et al. (2018) works well. Does this finding suggest that (1) is sufficient but not necessary in practice? More investigation is warranted. The effect of knockoff model misspecification on the false discovery rate (top panels) and power (bottom panels) when covariates are generated under a hidden Markov model (left panels) or a Markov chain model (right panels). Results with knockoffs generated under the correct model (dark or medium grey) are compared to results with knockoffs generated under the Gaussian approximation (light grey). The method for generating knockoffs has no noticeable effect on the false discovery rate or power. The parametric form of |$F_X$| is generally unknown, and therefore it is important to understand the consequences, in practice, of model misspecification in the generation of knockoffs. The work of Barber et al. (2018) provides an important first step in this direction. Code to reproduce Figs. 1 and 2 is available at https://github.com/jewellsean/gene_hunting_discussion. Sesia et al. (2019) were motivated by observing that in many practical settings, it is of interest to perform the conditional test of |$Y \perp\!\!\!\perp X_j \mid \{X_1,\ldots,X_p\} \setminus \{X_j\}$| rather than the marginal test of |$Y \perp\!\!\!\perp X_j$|⁠. However, in § 7 of their paper and § S7 of the supplementary material, the authors pre-processed the data by clustering and choosing a cluster representative. This drastically reduces the number of single nucleotide polymorphisms from 328 934 to 59 005 for the NFBC dataset and from 377 749 to 71 145 for the WTCCC dataset. The authors argue that pre-processing is required since neighbouring single nucleotide polymorphisms tend to be in high linkage disequilibrium and hence are highly correlated. Notwithstanding, this implies that the authors are in fact testing |$Y \perp\!\!\!\perp X_j \mid \{ X_{k} \}_{k \in \mathcal{A}\setminus j}$| instead of |$Y \perp\!\!\!\perp X_j \mid \{X_1,\ldots, X_p\}\setminus \{X_j\}$|⁠, where |$\mathcal{A}$| is the set of single nucleotide polymorphisms that survive the pre-processing procedure. Thus, the false discovery rate is ultimately controlled for clusters of single nucleotide polymorphisms, rather than for individual single nucleotide polymorphisms. Not all genetic variants are created equal, in the sense that some are more likely to be pathogenic, functional, or deleterious than others; see Kircher et al. (2014) and references therein. Practitioners may wish to incorporate hard-won prior knowledge in the testing process. The Bayesian variable importance statistics, introduced in § 4.1.2 of Candès et al. (2018), offer a potential way to exploit the prior information that is often present in genomic data while maintaining false discovery rate control. Although model-X knockoff procedures open the door to false discovery rate control in high-dimensional datasets, they come with the undesirable side effect that the set of selected features is random, since it depends on a realization of |$\tilde{X}$|⁠. Sesia et al. (2019) acknowledge the challenge of maintaining false discovery rate control while aggregating across multiple realizations of |$\tilde{X}$|⁠, but do not provide any guidance for the practitioner tasked with reporting discoveries. This is intellectually unsatisfying and poses problems in practice. We look forward to seeing new research on aggregation techniques that maintain control of the false discovery rate or another related measure. Jewell received funding from the Natural Sciences and Engineering Research Council of Canada. Witten was partially supported by a National Science Foundation CAREER Award, the National Institutes of Health, and a Simons Investigator Award in Mathematical Modeling of Living Systems.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: Expérimental (laboratoire)
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,119
Score d'incertitude au seuil0,281

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,016
Tête enseignante GPT0,269
Écart entre enseignants0,254 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2018
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueBiometrikaMême sujetGenetic Associations and EpidemiologyTravaux en français237 207