Phylogenomic synteny analysis tracks conserved ancient polyploid-derived triplicated genomic blocks across Asteraceae genomes
Bibliographic record
Abstract
Abstract The Asteraceae (Compositae) is the largest flowering plant family, ubiquitous in most terrestrial communities, and morphologically hyper-diverse. An ancient whole genome triplication (paleo-hexaploidization) occurred at approximately the same time as the evolutionary innovation and adaptive radiation of the family during the middle Eocene. Despite its importance, the genomic contents arising from this triplication have yet to be tracked in context of the Asteraceae genome evolution. We applied a synteny oriented phylogenomic analysis of 21 Asterales genomes and to study the paleo-hexaploidization and its consequences to gene, trait, and genome evolution. We identified 15 ancestral linkage groups (ALGs) that date back to the common diploid ancestor of all Asteraceae. Each of these groups was triplicated, resulting in 45 genomic blocks (3×15), which serve as the foundation for cross-family analyses. We demonstrate the complex evolutionary dynamics of the 45 genomic blocks across the Asteraceae phylogeny. We found that modern genomes are genetic mosaics of three progenitor genomes by extensive genomic exchange, chromosomal shuffling and gene fractionation. 157 genes retained three paleo-hexaploid derived syntenic paralogs across most Asteraceae species. Transcription factors (TFs) and auxin-related genes are significantly overrepresented in the conserved triplets, and expression of the paleo-hexaploidy paralogs is spatiotemporally differentiated. These genes are involved in the development of floral capitulum, a remarkable morphological innovation of the family. The discovery of conserved triplicated genes can direct further study to understand the evolutionary innovation, and the synteny-phylogenomic framework and ALGs provide a comparative framework to characterize newly sequenced Asteraceae genomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".