Discussion of ‘Gene hunting with hidden Markov model knockoffs’
Notice bibliographique
Résumé
We wish to congratulate the authors on their extension of model-X knockoffs beyond the Gaussian setting considered in Candès et al. (2018) to the well-studied class of Markov chain and hidden Markov models. This contribution builds upon a beautiful body of work that allows us to rethink false discovery rate control in the context of large and complex datasets (Barber & Candès, 2015, 2018; Weinstein et al., 2017; Barber et al., 2018; Candès et al., 2018; Katsevich & Sabatti, 2018). We believe that this innovation will help model-X knockoffs to be used in practical applications, such as genome-wide association studies explored in the paper, as well as other biological problems where Markov chain or hidden Markov models are common, see Sesia et al. (2019, §1.3) for some examples. In this discussion, we comment on two directions that merit further investigation. Sesia et al. (2019) provide an algorithm for generating model-X knockoffs without resorting to the Gaussian approximation of Candès et al. (2018), in the setting where the covariates are generated from a Markov chain or a hidden Markov model. This raises an intriguing question: if the covariates are distributed according to a Markov chain or hidden Markov model, then what could go wrong if we instead generate knockoffs according to the Gaussian approximation in Candès et al. (2018)? In practice, how bad would it be if we were to generate knockoffs the wrong way? Sesia et al. (2019) begin to address this issue within the context of the Crohn’s disease data in §7 of their paper, by comparing the set of variables selected when knockoffs are generated using a hidden Markov model with the set of variables selected when knockoffs are generated using the Gaussian approximation from Candès et al. (2018). They find that generating knockoffs under the hidden Markov model results in more discoveries than using the Gaussian approximation, and on this basis they claim that knockoffs generated under the hidden Markov model have higher power. However, ground truth data is not available for the Crohn’s disease data: the true positive single nucleotide polymorphisms are unknown. Therefore, we propose to study the false discovery rate and power on synthetic datasets. Covariates were generated under a Markov chain model, and knockoffs were generated using (a) the approach proposed in Sesia et al. (2019) and (b) the Gaussian approximation of Candès et al. (2018). Each panel displays the pairwise correlations of the covariates, |$\mathrm{cor}(X_{j}, X_{j-1})$|, versus the pairwise correlations of the knockoffs, |$\mathrm{cor}(\tilde{X}_{j}, \tilde{X}_{j-1})$|, for |$j=2,\ldots, p$|. The pairwise exchangeability condition (1) appears to be violated by the knockoffs generated using the Gaussian approximation. In practice, however, there is little difference between the knockoffs of Sesia et al. (2019) and those of Candès et al. (2018) in terms of power and false discovery rate control; see Fig. 2. In fact, the Gaussian approximation of Candès et al. (2018) works well. Does this finding suggest that (1) is sufficient but not necessary in practice? More investigation is warranted. The effect of knockoff model misspecification on the false discovery rate (top panels) and power (bottom panels) when covariates are generated under a hidden Markov model (left panels) or a Markov chain model (right panels). Results with knockoffs generated under the correct model (dark or medium grey) are compared to results with knockoffs generated under the Gaussian approximation (light grey). The method for generating knockoffs has no noticeable effect on the false discovery rate or power. The parametric form of |$F_X$| is generally unknown, and therefore it is important to understand the consequences, in practice, of model misspecification in the generation of knockoffs. The work of Barber et al. (2018) provides an important first step in this direction. Code to reproduce Figs. 1 and 2 is available at https://github.com/jewellsean/gene_hunting_discussion. Sesia et al. (2019) were motivated by observing that in many practical settings, it is of interest to perform the conditional test of |$Y \perp\!\!\!\perp X_j \mid \{X_1,\ldots,X_p\} \setminus \{X_j\}$| rather than the marginal test of |$Y \perp\!\!\!\perp X_j$|. However, in § 7 of their paper and § S7 of the supplementary material, the authors pre-processed the data by clustering and choosing a cluster representative. This drastically reduces the number of single nucleotide polymorphisms from 328 934 to 59 005 for the NFBC dataset and from 377 749 to 71 145 for the WTCCC dataset. The authors argue that pre-processing is required since neighbouring single nucleotide polymorphisms tend to be in high linkage disequilibrium and hence are highly correlated. Notwithstanding, this implies that the authors are in fact testing |$Y \perp\!\!\!\perp X_j \mid \{ X_{k} \}_{k \in \mathcal{A}\setminus j}$| instead of |$Y \perp\!\!\!\perp X_j \mid \{X_1,\ldots, X_p\}\setminus \{X_j\}$|, where |$\mathcal{A}$| is the set of single nucleotide polymorphisms that survive the pre-processing procedure. Thus, the false discovery rate is ultimately controlled for clusters of single nucleotide polymorphisms, rather than for individual single nucleotide polymorphisms. Not all genetic variants are created equal, in the sense that some are more likely to be pathogenic, functional, or deleterious than others; see Kircher et al. (2014) and references therein. Practitioners may wish to incorporate hard-won prior knowledge in the testing process. The Bayesian variable importance statistics, introduced in § 4.1.2 of Candès et al. (2018), offer a potential way to exploit the prior information that is often present in genomic data while maintaining false discovery rate control. Although model-X knockoff procedures open the door to false discovery rate control in high-dimensional datasets, they come with the undesirable side effect that the set of selected features is random, since it depends on a realization of |$\tilde{X}$|. Sesia et al. (2019) acknowledge the challenge of maintaining false discovery rate control while aggregating across multiple realizations of |$\tilde{X}$|, but do not provide any guidance for the practitioner tasked with reporting discoveries. This is intellectually unsatisfying and poses problems in practice. We look forward to seeing new research on aggregation techniques that maintain control of the false discovery rate or another related measure. Jewell received funding from the Natural Sciences and Engineering Research Council of Canada. Witten was partially supported by a National Science Foundation CAREER Award, the National Institutes of Health, and a Simons Investigator Award in Mathematical Modeling of Living Systems.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».