Haplotype‐phased and chromosome‐level genome assembly of <i>Puccinia polysora</i>, a giga‐scale fungal pathogen causing southern corn rust
Bibliographic record
Abstract
Rust fungi are characterized by large genomes with high repeat content and have two haploid nuclei in most life stages, which makes achieving high-quality genome assemblies challenging. Here, we described a pipeline using HiFi reads and Hi-C data to assemble a gigabase-sized fungal pathogen, Puccinia polysora f.sp. zeae, to haplotype-phased and chromosome-scale. The final assembled genome is 1.71 Gbp, with ~850 Mbp and 18 chromosomes in each haplotype, being currently one of the two giga-scale fungi assembled to chromosome level. Transcript-based annotation identified 47,512 genes for the dikaryotic genome with a similar number for each haplotype. A high level of interhaplotype variation was found with 10% haplotype-specific BUSCO genes, 5.8 SNPs/kbp, and structural variation accounting for 3% of the genome size. The P. polysora genome displayed over 85% repeat contents, with genome-size expansion and copy number increasing of species-specific orthogroups. Interestingly, these features did not affect overall synteny with other Puccinia species having smaller genomes. Fine-time-point transcriptomics revealed seven clusters of coexpressed secreted proteins that are conserved between two haplotypes. The fact that candidate effectors interspersed with all genes indicated the absence of a "two-speed genome" evolution in P. polysora. Genome resequencing of 79 additional isolates revealed a clonal population structure of P. polysora in China with low geographic differentiation. Nevertheless, a minor population differentiated from the major population by having mutations on secreted proteins including AvrRppC, indicating the ongoing virulence to evade recognition by RppC, a major resistance gene in Chinese corn cultivars. The high-quality assembly provides valuable genomic resources for future studies on disease management and the evolution of P. polysora.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".