Whole genome duplication: challenges and considerations associated with sequence orthology assignment in Salmoninae
Bibliographic record
Abstract
To illustrate some of the challenges and considerations in assigning correct orthology necessary for any comparative genomic investigation among salmonids, sequence data from the non-coding regions of different chromosomes in three members of the subfamily Salmoninae, rainbow trout Oncorhynchus mykiss, Atlantic salmon Salmo salar and Arctic charr Salvelinus alpinus, were compared. By analysing c. 55 distinct loci, corresponding to c. 142 kbp sequence information per species, 18 duplicated patterns representative of the two sequential rounds of teleost-specific whole genome duplications (i.e. 3R and 4R WGD) were identified. Sequence similarities between the 4R paralogues were c. 90%, which was slightly lower than those of the 4R orthologues and c. 60% for the 3R products. Through careful examination of the sequence data, however, only 14 loci could reliably be assigned as true orthologues. Locus-specific trees were constructed through maximum parsimony, maximum likelihood and neighbour-joining methods and were rooted using the information from a close relative, lake whitefish Coregonus clupeaformis. All approaches generated congruent trees supporting the {Coregonus [Salmo (Oncorhynchus, Salvelinus)]} topology. The general phenotypic characteristics of sequences, however, were highly suggestive of the basal position of Oncorhynchus, raising the hypothesis of an accelerated rate of nucleotide evolution in this species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.021 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".