UPGRADING THE DURUM WHEAT GENOMIC RESOURCES: FROM THEPLATINUM-QUALITY SVEVO GENOME
Bibliographic record
Abstract
Durum wheat (Triticum turgidum L. ssp. durum) is a major cereal and staple in the semi-arid regions of the Mediterranean Basin. It originates from BBAA wild tetraploid domesticated in Neolithic era, later evolving to domesticated emmer and then to up to 11 T. turgidum subspecies, including durum wheat landraces and modern cultivars. Tetraploid wheat is the donor of the A and B genomes of hexaploid bread wheat (DDAABB), representing therefore a valuable source of genetic variability and beneficial alleles for both durum and bread wheat breeding. The reference durum wheat genome (cv Svevo) was previously sequenced and assembled using a short-read sequencing approach (Maccaferri et al., 2009), representing an important milestone for wheat genomics. In an effort to improve this release, an international consortium was established to produce a platinum-quality reference genome, that accomplish with the Contiguity, Completeness, Correctness ultimate RefSeq2.0 requirements. PACBIO HiFi long read 35X sequencing was coupled to BIONANO Optical Mapping, producing 259 Hybrid scaffolds (N50 = 112.3 Mb) ordered by Hi-C data in a 10.4 Gb assembly. A complete and accurate gene prediction was then obtained by coupling Illumina RNASeq and Nanopore Isoseq sequencing of 30 plant samples representative of a range of tissues and developmental stages grown in normal growth condition and of 28 samples obtained under biotic, abiotic and nutrient stress conditions. The expression of the 68,154 high confidence genes will be integrated in transcriptome atlas, together with more than 100,000 low confidence, TE-related or long non-coding genes. Gene regulatory networks associated to spike and seed development are under investigation and will be associate with the chromatin accessibility obtained by ATAC-Seq assays from the same samples, providing an exceptional starting point to study how enhancers, promoters, transcription factors binding cooperatively regulate gene expression during development. Finally, the preliminary results coming from the sequencing of 24 tetraploid wheat (representative of wild emmer, domesticated emmer, turgidum and turanicum subspecies, durum wheat landraces and modern varieties) pangenome will be presented. Genomic and transcriptomic data coming from this project will allow to fully explore the huge tetraploid genome diversity available for the identification of beneficial alleles enhancing both durum and bread wheat resilience to abiotic stresses a disease resistance. This research has been supported by: PRIN project “PanWheatGrain - Grain Pangenomics for Durum Wheat Sustainable Production” funded by the Italian Ministry of University Education and Research; Agritech National Research Center funded by the European Union Next-Generation-EU (PNRR - MISSIONE 4 COMPONENTE 2, INVESTIMENTO 1.4 – D.D. 1032 17/06/2022, CN00000022) and by the Canadian project “Canadian Tetraploid Pan Genomics Project”.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".