Diversification of Hox Gene Clusters in Osteoglossomorph Fish in Comparison to Other Teleosts and the Spotted Gar Outgroup
Bibliographic record
Abstract
An ancient genome duplication (TGD or 3R) occurred in teleost fish after divergence from the lineage leading to gar. This genome duplication is shared by the three extant teleost lineages: Osteoglossomorpha (bony-tongues), Elopomorpha (eels and tarpons), and Clupeocephala (a large clade including salmon, carp, medaka, zebrafish, cichlids, pufferfish, stickleback, and ∼26,000 other species). After TGD, different clupeocephalan species retained different gene duplicates; this is seen clearly in Hox gene clusters but extends to all genes. Since divergent resolution of TGD paralogs is a potential driving force for speciation, it is possible this contributed to diversification of this clade. The extent to which divergent resolution of TGD paralogs occurred within Osteoglossomorpha has not been investigated in detail, and Hox cluster organization has been reported for just two species: Pantodon buchholzi (Pantodontidae) and Scleropages formosus (Osteoglossidae). We applied survey-scale genome sequencing and de novo assembly to three further osteoglossomorph taxa: Osteoglossum bicirrhosum (Osteoglossidae), Chitala ornata (Notopteridae), and Gnathonemus petersii (Mormyridae). We find that each retained more Hox genes than clupeocephalan taxa (excluding those that underwent additional genome duplication), but fewer than eels. Several Hox genes are missing in all teleosts, including duplicates of two Hox genes present in the slow evolving pre-TGD genome of the spotted gar. We find divergent resolution through individual gene losses, and whole cluster losses have been rampant across osteoglossomorphs, despite their extant species paucity. We suggest that reciprocal gene loss following TGD was probably insufficient to drive the exceptional diversification of teleosts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".