Multivariate Genome-Wide Association Analyses Reveal the Genetic Basis of Seed Fatty Acid Composition in Oat ( <i>Avena sativa</i> L.)
Bibliographic record
Abstract
L.) has a high concentration of oils, comprised primarily of healthful unsaturated oleic and linoleic fatty acids. To accelerate oat plant breeding efforts, we sought to identify loci associated with variation in fatty acid composition, defined as the types and quantities of fatty acids. We genotyped a panel of 500 oat cultivars with genotyping-by-sequencing and measured the concentrations of ten fatty acids in these oat cultivars grown in two environments. Measurements of individual fatty acids were highly correlated across samples, consistent with fatty acids participating in shared biosynthetic pathways. We leveraged these phenotypic correlations in two multivariate genome-wide association study (GWAS) approaches. In the first analysis, we fitted a multivariate linear mixed model for all ten fatty acids simultaneously while accounting for population structure and relatedness among cultivars. In the second, we performed a univariate association test for each principal component (PC) derived from a singular value decomposition of the phenotypic data matrix. To aid interpretation of results from the multivariate analyses, we also conducted univariate association tests for each trait. The multivariate mixed model approach yielded 148 genome-wide significant single-nucleotide polymorphisms (SNPs) at a 10% false-discovery rate, compared to 129 and 73 significant SNPs in the PC and univariate analyses, respectively. Thus, explicit modeling of the correlation structure between fatty acids in a multivariate framework enabled identification of loci associated with variation in seed fatty acid concentration that were not detected in the univariate analyses. Ultimately, a detailed characterization of the loci underlying fatty acid variation can be used to enhance the nutritional profile of oats through breeding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".