Construction and Analysis of a Rice Pan-genome Reveals Structural Variation Hotspots Across Subspecies
Bibliographic record
Abstract
Rice ( Oryza sativa ) is a staple cereal with immense global importance, yet a single reference genome cannot capture the full genetic diversity underlying key traits. Pan-genomics has emerged as a paradigm to characterize the “pan-genome”-the total genomic repertoire of a species-including core genes shared by all accessions and dispensable genes present in some but absent in others. Here, this study reviews the construction and analysis of rice pan-genomes and the insights they provide into structural variation hotspots across rice subspecies; outlines how the limitations of a single reference genome have driven the development of plant pan-genomics, enabling discovery of extensive genomic variation that was previously hidden; describes strategies for building rice pan-genomes, from early short-read sequencing approaches to recent long-read assemblies and graph-based genome models that integrate diverse accessions ( indica , japonica , aus , aromatic , wild relatives). Major types of structural variation-insertions, deletions, inversions, translocations, copy number variations-are defined, and this study surveys computational tools for their detection; synthesizes findings on the distribution of structural variants (SVs) in the rice genome and identify hotspots of variation specific to certain lineages. The functional impact of SVs is discussed, with case studies linking structural variants to agronomic traits (yield, stress tolerance, flowering time) and to gene presence/absence variation affecting gene families (e.g. disease resistance genes). Comparative pan-genome analyses across rice subspecies illuminate how evolutionary forces like domestication bottlenecks, introgression, and selection have shaped genomic differences between indica and japonica rice. Finally, this study highlights emerging applications of rice pan-genome research in germplasm utilization, genome-wide association studies, marker-assisted breeding, and de novo domestication, and discusses future prospects and challenges in integrating multi-omics data and developing pan-genomic resources for sustainable agriculture under climate change.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".