Genetic Variation and Population Structure of Oryza glaberrima and Development of a Mini-Core Collection Using DArTseq
Bibliographic record
Abstract
The sequence variation present in accessions conserved in gene banks can best be used in plant improvement when it is properly characterized and published. Using low cost and high density single nucleotide polymorphism (SNP) assays, the genetic diversity, population structure, and relatedness between pairs of accessions can be quickly assessed. This information is relevant for different purposes, including creating core and mini-core sets that represent the maximum possible genetic variation contained in the whole collection. Here, we studied the genetic variation and population structure of 2,179 Oryza glaberrima Steud. accessions conserved at the AfricaRice gene bank using 27,560 DArTseq-based SNPs. Only 14% (3,834 of 27,560) of the SNPs were polymorphic across the 2,179 accessions, which is much lower than diversity reported in other Oryza species. Genetic distance between pairs of accessions varied from 0.005 to 0.306, with 1.5% of the pairs nearly identical, 8.0% of the pairs similar, 78.1% of the pairs moderately distant, and 12.5% of the pairs very distant. The number of redundant accessions that contribute little or no new genetic variation to the O. glaberrima collection was very low. Using the maximum length sub-tree method, we propose a subset of 1,330 and 350 accessions to represent a core and mini-core collection, respectively. The core and mini-core sets accounted for ~ 61 and 16%, respectively, of the whole collection, and captured 97-99% of the SNP polymorphism and nearly all allele and genotype frequencies observed in the whole O. glaberrima collection available at the AfricaRice gene bank. Cluster, principal component and model-based population structure analyses all divided the 2,179 accessions into five groups, based roughly on country of origin but less so on ecology. The first, third and fourth groups consisted of accessions primarily from Liberia, Nigeria, and Mali, respectively; the second group consisted primarily of accessions from Togo and Nigeria; and the fifth and smallest group was a mixture of accessions from multiple countries. Analysis of molecular variance showed between 10.8- 28.9% of the variation among groups with the remaining 71.1-89.2% attributable to differences within groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".