Analysis of Intraspecies Diversity in Wheat and Barley Genomes Identifies Breakpoints of Ancient Haplotypes and Provides Insight into the Structure of Diploid and Hexaploid Triticeae Gene Pools
Bibliographic record
Abstract
A large number of wheat (Triticum aestivum) and barley (Hordeum vulgare) varieties have evolved in agricultural ecosystems since domestication. Because of the large, repetitive genomes of these Triticeae crops, sequence information is limited and molecular differences between modern varieties are poorly understood. To study intraspecies genomic diversity, we compared large genomic sequences at the Lr34 locus of the wheat varieties Chinese Spring, Renan, and Glenlea, and diploid wheat Aegilops tauschii. Additionally, we compared the barley loci Vrs1 and Rym4 of the varieties Morex, Cebada Capa, and Haruna Nijo. Molecular dating showed that the wheat D genome haplotypes diverged only a few thousand years ago, while some barley and Ae. tauschii haplotypes diverged more than 500,000 years ago. This suggests gene flow from wild barley relatives after domestication, whereas this was rare or absent in the D genome of hexaploid wheat. In some segments, the compared haplotypes were very similar to each other, but for two varieties each at the Rym4 and Lr34 loci, sequence conservation showed a breakpoint that separates a highly conserved from a less conserved segment. We interpret this as recombination breakpoints of two ancient haplotypes, indicating that the Triticeae genomes are a heterogeneous and variable mosaic of haplotype fragments. Analysis of insertions and deletions showed that large events caused by transposable element insertions, illegitimate recombination, or unequal crossing over were relatively rare. Most insertions and deletions were small and caused by template slippage in short homopolymers of only a few base pairs in size. Such frequent polymorphisms could be exploited for future molecular marker development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".