The role of common genetic variation in model polygenic and monogenic traits
Bibliographic record
Abstract
The aim of this thesis is to explore the role of common genetic variation, identified through genome-wide association (GWA) studies, in human traits and diseases, using height as a model polygenic trait, type 2 diabetes as a model common polygenic disease, and maturity onset diabetes of the young (MODY) as a model monogenic disease. The wave of the initial GWA studies, such as the Wellcome Trust Case-Control Consortium (WTCCC) study of seven common diseases, substantially increased the number of common variants associated with a range of different multifactorial traits and diseases. The initial excitement, however, seems to have been followed by some disappointment that the identified variants explain a relatively small proportion of the genetic variance of the studied trait, and that only few large effect or causal variants have been identified. Inevitably, this has led to criticism of the GWA studies, mainly that the findings are of limited clinical, or indeed scientific, benefit. Using height as a model, Chapter 2 explores the utility of GWA studies in terms of identifying regions that contain relevant genes, and in answering some general questions about the genetic architecture of highly polygenic traits. Chapter 3 takes this further into a large collaborative study and the largest sample size in a GWA study to date, mainly focusing on demonstrating the biological relevance of the identified variants, even when a large number of associated regions throughout the genome is implicated by these associations. Furthermore, it shows examples of different features of the genetic architecture, such as allelic heterogeneity and pleiotropy. Chapter 4 looks at the predictive value and, therefore, clinical utility, of variants found to associate with type 2 diabetes, a common multifactorial disease that is increasing in prevalence despite known environmental risk factors. This is a disease where knowledge of the genetic risk has potentially substantial clinical relevance. Finally, Chapter 5 approaches the monogenic-polygenic disease bridge in the direction opposite to that approached in the past: most studies have investigated genes mutated in monogenic diseases as candidates for harboring common variants predisposing to related polygenic diseases. This chapter looks at the common type 2 diabetes variants as modifiers of disease onset in patients with a monogenic but clinically heterogeneous disease, maturity onset diabetes of the young (MODY).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".