Phenotyping and Latent Genetic Interaction Analysis using Quantile Regression
Notice bibliographique
Résumé
Linear regression is commonly used in genome-wide association studies (GWAS) with continuous outcomes and examines the conditional mean of a trait assuming normality. The real possibility of model misspecification and the scientific interest beyond the conditional mean makes quantile regression, which explores the conditional distribution of a quantitative trait without the assumption of normality, an attractive alternative framework, with application in genetics and genomics growing in recent years. In this thesis, I demonstrate the utility of quantile regression for genetic studies motivated by challenges identifying genetic contributions to variation in cystic fibrosis (CF) disease severity. We implemented quantile regression (1) to build a lung disease phenotype for Canadians with CF, to be used for genetic association studies; and (2) to develop a powerful and robust association test that leverages latent genetic interactions. For (1), I derived Canadian CF-specific forced expiratory volume in 1 second (FEV1) reference equations based on the Canadian CF registry which captures the clinical experience of the Canadian CF population. These CF-specific FEV1 reference equations were the building blocks of the lung disease phenotype in the analysis of genetic association of the Canadian CF Gene Modifier Study. For (2), I provide a review of association tests that leverage latent genetic interactions in the literature, namely joint location and scale tests which rely on a normality assumption, and perform simulation studies to investigate the methodological strengths and weaknesses. To address the limitations of existing methods identified via simulation, I propose a new joint location and scale test based on quantile regression (qJLS) that is free of distributional assumptions, thus applies to non-Gaussian traits. The qJLS is as powerful as the existing joint location and scale tests under Gaussian traits and is computationally efficient for GWAS. Simulation studies evaluated the properties of the qJLS and its performance compared to existing approaches. Application of the qJLS to a GWAS of CF lung disease in the Canadian CF Gene Modifier Study identified novel putatively contributing loci and demonstrates it as a powerful alternative to conventional genetic association tests, where interactions may contribute to a quantitative trait.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,043 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,002 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,002 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».