A Highly Specific Genome-Wide Association Study Integrated with Transcriptome Data Reveals the Contribution of Copy Number Variations to Specialized Metabolites in Arabidopsis thaliana Accessions
Bibliographic record
Abstract
Lineage-specific gene duplications contribute to a large variation in specialized metabolites among different plant species. There is also considerable variability in the specialized metabolites within a single plant species. However, it is unclear whether copy number variations (CNVs) derived from gene duplication events contribute to the diversity of specialized metabolites within species. We identified metabolome quantitative trait genes (mQTGs) associated with quantitative metabolite variations and examined the relationship between mQTGs and CNVs. We obtained 1,335 specialized metabolite signals from 53 worldwide A. thaliana accessions using liquid chromatography-quadrupole time-of-flight mass spectrometry. In this study, genes associated with specialized metabolites were inferred by either a generally authorized genome-wide association study (GWAS) approach or a novel analysis of the association between gene expression and metabolite accumulation. Genes qualified by both analyses are defined to be mQTGs. The integrated method enabled us to detect mQTGs with a low false positive rate (=5.71 × 10-4). We also identified 5,654 genes associated with 1,335 specialized metabolites. Of these genes, 4.4% were affected by CNVs, which was more than expected (χ2 test: P < 0.01). This result suggests that CNVs contribute to variations in specialized metabolites within a species. To assess the contribution of CNVs to adaptive evolution in A. thaliana, we examined the selective sweeps around the mQTGs. We observed that the mQTGs with CNVs tended to undergo selective sweeps. These observations imply that variations in specialized metabolites caused by CNVs contribute to the adaptive evolution of A. thaliana.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".