The role of common genetic variation in model polygenic and monogenic traits
Bibliographic record
Abstract
The aim of this thesis is to explore the role of common genetic variation, identified through genome-wide association (GWA) studies, in human traits and diseases, using height as a model polygenic trait, type 2 diabetes as a model common polygenic disease, and maturity onset diabetes of the young (MODY) as a model monogenic disease. \nThe wave of the initial GWA studies, such as the Wellcome Trust Case-Control Consortium (WTCCC) study of seven common diseases, substantially increased the number of common variants associated with a range of different multifactorial traits and diseases. The initial excitement, however, seems to have been followed by some disappointment that the identified variants explain a relatively small proportion of the genetic variance of the studied trait, and that only few large effect or causal variants have been identified. Inevitably, this has led to criticism of the GWA studies, mainly that the findings are of limited clinical, or indeed scientific, benefit. \nUsing height as a model, Chapter 2 explores the utility of GWA studies in terms of identifying regions that contain relevant genes, and in answering some general questions about the genetic architecture of highly polygenic traits. \nChapter 3 takes this further into a large collaborative study and the largest sample size in a GWA study to date, mainly focusing on demonstrating the biological relevance of the identified variants, even when a large number of associated regions throughout the genome is implicated by these associations. Furthermore, it shows examples of different features of the genetic architecture, such as allelic heterogeneity and pleiotropy. \nChapter 4 looks at the predictive value and, therefore, clinical utility, of variants found to associate with type 2 diabetes, a common multifactorial disease that is increasing in prevalence despite known environmental risk factors. This is a disease where knowledge of the genetic risk has potentially substantial clinical relevance. \nFinally, Chapter 5 approaches the monogenic-polygenic disease bridge in the direction opposite to that approached in the past: most studies have investigated genes mutated in monogenic diseases as candidates for harboring common variants predisposing to related polygenic diseases. This chapter looks at the common type 2 diabetes variants as modifiers of disease onset in patients with a monogenic but clinically heterogeneous disease, maturity onset diabetes of the young (MODY).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".