Genetically Distinct Rice Lines for Specific Characters as Revealed by Gene-Associated Average Pairwise Dissimilarity
Bibliographic record
Abstract
Broadening the genetic base of an elite breeding gene pool is one important goal in a successful long-term plant breeding program. This goal is largely achieved through the search for and introgression of exotic germplasm with adaptive traits. However, little is known about the genetic backgrounds of acquired exotic germplasm, as germplasm selection is mainly based on trait information. Here, we expanded an average pairwise dissimilarity (APD) analysis to samples with SNP genotypes associated with genes for specific characters of breeding interest. Specifically, we explored a gene-associated APD analysis in a genomic characterization of 2643 rice lines based on their published FASTQ data. Published contigs for cloned genes conditioning heat tolerance, cold tolerance, fertility, and seed size were downloaded as gene reference sequences for SNP calling, along with those SNP calls based on the rice reference genome and published indels. Totally, eight SNP or indel data sets were formed for each of three sample groups (All2643, Indica1789, and Japonica854). APD estimation was made for each of the 24 data sets. For each sample group, four novel sets of the 25 most genetically distinct rice lines, each for an assayed character, were identified. Further analyses of APD estimates also revealed some interesting APD properties. Four contig-based SNP data sets for four specific characters displayed similar APD frequency distributions and positive high correlations of APD estimates. Contig-based APD estimates were negatively correlated with genome-based APD estimates and nearly uncorrelated with indel-based APD estimates. These findings are significant for plant germplasm characterization and germplasm utilization in plant breeding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".