MétaCan
Menu
Back to cohort
Record W2914988643 · doi:10.1093/biomet/asy061

Discussion of ‘Gene hunting with hidden Markov model knockoffs’

2018· article· en· W2914988643 on OpenAlexfundaboutno aff
Sean Jewell, Daniela Witten

Bibliographic record

VenueBiometrika · 2018
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Associations and Epidemiology
Canadian institutionsnot available
FundersNational Institute of Biomedical Imaging and BioengineeringNational Institute of General Medical SciencesNational Institute on Drug AbuseNatural Sciences and Engineering Research Council of CanadaNational Institutes of HealthNational Science Foundation
KeywordsHidden Markov modelMathematicsMarkov chainMarkov modelArtificial intelligenceStatisticsComputer science

Abstract

fetched live from OpenAlex

We wish to congratulate the authors on their extension of model-X knockoffs beyond the Gaussian setting considered in Candès et al. (2018) to the well-studied class of Markov chain and hidden Markov models. This contribution builds upon a beautiful body of work that allows us to rethink false discovery rate control in the context of large and complex datasets (Barber & Candès, 2015, 2018; Weinstein et al., 2017; Barber et al., 2018; Candès et al., 2018; Katsevich & Sabatti, 2018). We believe that this innovation will help model-X knockoffs to be used in practical applications, such as genome-wide association studies explored in the paper, as well as other biological problems where Markov chain or hidden Markov models are common, see Sesia et al. (2019, §1.3) for some examples. In this discussion, we comment on two directions that merit further investigation. Sesia et al. (2019) provide an algorithm for generating model-X knockoffs without resorting to the Gaussian approximation of Candès et al. (2018), in the setting where the covariates are generated from a Markov chain or a hidden Markov model. This raises an intriguing question: if the covariates are distributed according to a Markov chain or hidden Markov model, then what could go wrong if we instead generate knockoffs according to the Gaussian approximation in Candès et al. (2018)? In practice, how bad would it be if we were to generate knockoffs the wrong way? Sesia et al. (2019) begin to address this issue within the context of the Crohn’s disease data in §7 of their paper, by comparing the set of variables selected when knockoffs are generated using a hidden Markov model with the set of variables selected when knockoffs are generated using the Gaussian approximation from Candès et al. (2018). They find that generating knockoffs under the hidden Markov model results in more discoveries than using the Gaussian approximation, and on this basis they claim that knockoffs generated under the hidden Markov model have higher power. However, ground truth data is not available for the Crohn’s disease data: the true positive single nucleotide polymorphisms are unknown. Therefore, we propose to study the false discovery rate and power on synthetic datasets. Covariates were generated under a Markov chain model, and knockoffs were generated using (a) the approach proposed in Sesia et al. (2019) and (b) the Gaussian approximation of Candès et al. (2018). Each panel displays the pairwise correlations of the covariates, |$\mathrm{cor}(X_{j}, X_{j-1})$|⁠, versus the pairwise correlations of the knockoffs, |$\mathrm{cor}(\tilde{X}_{j}, \tilde{X}_{j-1})$|⁠, for |$j=2,\ldots, p$|⁠. The pairwise exchangeability condition (1) appears to be violated by the knockoffs generated using the Gaussian approximation. In practice, however, there is little difference between the knockoffs of Sesia et al. (2019) and those of Candès et al. (2018) in terms of power and false discovery rate control; see Fig. 2. In fact, the Gaussian approximation of Candès et al. (2018) works well. Does this finding suggest that (1) is sufficient but not necessary in practice? More investigation is warranted. The effect of knockoff model misspecification on the false discovery rate (top panels) and power (bottom panels) when covariates are generated under a hidden Markov model (left panels) or a Markov chain model (right panels). Results with knockoffs generated under the correct model (dark or medium grey) are compared to results with knockoffs generated under the Gaussian approximation (light grey). The method for generating knockoffs has no noticeable effect on the false discovery rate or power. The parametric form of |$F_X$| is generally unknown, and therefore it is important to understand the consequences, in practice, of model misspecification in the generation of knockoffs. The work of Barber et al. (2018) provides an important first step in this direction. Code to reproduce Figs. 1 and 2 is available at https://github.com/jewellsean/gene_hunting_discussion. Sesia et al. (2019) were motivated by observing that in many practical settings, it is of interest to perform the conditional test of |$Y \perp\!\!\!\perp X_j \mid \{X_1,\ldots,X_p\} \setminus \{X_j\}$| rather than the marginal test of |$Y \perp\!\!\!\perp X_j$|⁠. However, in § 7 of their paper and § S7 of the supplementary material, the authors pre-processed the data by clustering and choosing a cluster representative. This drastically reduces the number of single nucleotide polymorphisms from 328 934 to 59 005 for the NFBC dataset and from 377 749 to 71 145 for the WTCCC dataset. The authors argue that pre-processing is required since neighbouring single nucleotide polymorphisms tend to be in high linkage disequilibrium and hence are highly correlated. Notwithstanding, this implies that the authors are in fact testing |$Y \perp\!\!\!\perp X_j \mid \{ X_{k} \}_{k \in \mathcal{A}\setminus j}$| instead of |$Y \perp\!\!\!\perp X_j \mid \{X_1,\ldots, X_p\}\setminus \{X_j\}$|⁠, where |$\mathcal{A}$| is the set of single nucleotide polymorphisms that survive the pre-processing procedure. Thus, the false discovery rate is ultimately controlled for clusters of single nucleotide polymorphisms, rather than for individual single nucleotide polymorphisms. Not all genetic variants are created equal, in the sense that some are more likely to be pathogenic, functional, or deleterious than others; see Kircher et al. (2014) and references therein. Practitioners may wish to incorporate hard-won prior knowledge in the testing process. The Bayesian variable importance statistics, introduced in § 4.1.2 of Candès et al. (2018), offer a potential way to exploit the prior information that is often present in genomic data while maintaining false discovery rate control. Although model-X knockoff procedures open the door to false discovery rate control in high-dimensional datasets, they come with the undesirable side effect that the set of selected features is random, since it depends on a realization of |$\tilde{X}$|⁠. Sesia et al. (2019) acknowledge the challenge of maintaining false discovery rate control while aggregating across multiple realizations of |$\tilde{X}$|⁠, but do not provide any guidance for the practitioner tasked with reporting discoveries. This is intellectually unsatisfying and poses problems in practice. We look forward to seeing new research on aggregation techniques that maintain control of the false discovery rate or another related measure. Jewell received funding from the Natural Sciences and Engineering Research Council of Canada. Witten was partially supported by a National Science Foundation CAREER Award, the National Institutes of Health, and a Simons Investigator Award in Mathematical Modeling of Living Systems.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.119
Threshold uncertainty score0.281

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.016
GPT teacher head0.269
Teacher spread0.254 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2018
Admission routes2
Has abstractyes

Explore more

Same venueBiometrikaSame topicGenetic Associations and EpidemiologyFrench-language works237,207