Phenotyping and Latent Genetic Interaction Analysis using Quantile Regression
Bibliographic record
Abstract
Linear regression is commonly used in genome-wide association studies (GWAS) with continuous outcomes and examines the conditional mean of a trait assuming normality. The real possibility of model misspecification and the scientific interest beyond the conditional mean makes quantile regression, which explores the conditional distribution of a quantitative trait without the assumption of normality, an attractive alternative framework, with application in genetics and genomics growing in recent years. In this thesis, I demonstrate the utility of quantile regression for genetic studies motivated by challenges identifying genetic contributions to variation in cystic fibrosis (CF) disease severity. We implemented quantile regression (1) to build a lung disease phenotype for Canadians with CF, to be used for genetic association studies; and (2) to develop a powerful and robust association test that leverages latent genetic interactions. For (1), I derived Canadian CF-specific forced expiratory volume in 1 second (FEV1) reference equations based on the Canadian CF registry which captures the clinical experience of the Canadian CF population. These CF-specific FEV1 reference equations were the building blocks of the lung disease phenotype in the analysis of genetic association of the Canadian CF Gene Modifier Study. For (2), I provide a review of association tests that leverage latent genetic interactions in the literature, namely joint location and scale tests which rely on a normality assumption, and perform simulation studies to investigate the methodological strengths and weaknesses. To address the limitations of existing methods identified via simulation, I propose a new joint location and scale test based on quantile regression (qJLS) that is free of distributional assumptions, thus applies to non-Gaussian traits. The qJLS is as powerful as the existing joint location and scale tests under Gaussian traits and is computationally efficient for GWAS. Simulation studies evaluated the properties of the qJLS and its performance compared to existing approaches. Application of the qJLS to a GWAS of CF lung disease in the Canadian CF Gene Modifier Study identified novel putatively contributing loci and demonstrates it as a powerful alternative to conventional genetic association tests, where interactions may contribute to a quantitative trait.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.043 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".