Reference genome bias in light of species-specific chromosomal reorganization and translocations
Bibliographic record
Abstract
Abstract Background Whole-genome sequencing efforts, have during the past decade, unveiled the central role of genomic rearrangements—such as chromosomal inversions—in evolutionary processes, including local adaptation in a wide range of taxa. However, employment of reference genomes from distantly or even closely related species for mapping and the subsequent variant calling can lead to errors and/or biases in the datasets generated for downstream analyses. Results Here, we capitalize on the recently generated chromosome-anchored genome assemblies for Arctic cod ( Arctogadus glacialis ), polar cod ( Boreogadus saida ), and Atlantic cod ( Gadus morhua ) to evaluate the extent and consequences of reference bias on population sequencing datasets (approx. 15–20 × coverage) for both Arctic cod and polar cod. Our findings demonstrate that the choice of reference genome impacts the mapping statistics, including mapping depth and mapping quality, as well as core population genetic estimates, such as heterozygosity levels, nucleotide diversity (π), and cross-species genetic divergence (D XY ). Furthermore, using a more distantly related reference genome can lead to inaccurate detection and characterization of chromosomal inversions, i.e., in terms of size (length) and location (position), due to inter-chromosomal reorganizations between species. Additionally, we observe that some of the verified species-specific inversions are split across multiple genomic regions when mapped against a heterospecific reference. Conclusions Inaccurate identification of chromosomal rearrangements as well as biased population genetic measures could potentially lead to erroneous interpretation of species-specific genomic diversity, impede the resolution of local adaptation, and thus, impact predictions of their genomic potential to respond to climatic and other environmental perturbations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".