A Beginner’s Guide to Structural Variants in Eco-Evolutionary Population Genomics: Everything You Wanted to Know
Notice bibliographique
Résumé
1. IntroductionCharacterizing genomic variation is fundamental to address many ecological and evolutionary questions. The ‘evolution’ of DNA sequencing methods over time has enabled the identification of new aspects of genetic diversity at every step. In particular, ecological and evolutionary genomics have flourished with the increased availability of high quality reference genomes (Formenti et al., 2022), and attainability of whole genome sequencing (WGS) (Fuentes-Pardo & Ruzzante, 2017). Resequencing entire genomes, rather than a small portion through reduced-representation approaches, provides a rich source of information and has led to a proliferation of methods to investigate evolutionary and demographic processes. We can now identify signatures of balancing selection in the genome (Stern & Lee, 2020), reconstruct demographic history in the near and distant past with unprecedented resolution (Nadachowska-Brzyska et al., 2015; Santiago et al., 2020), and characterize the roles of genome structure and recombination in the levels and distribution of genomic variation across the genome (Akopyan et al., 2025; Tigano et al., 2021). Many genomic analysis methods are currently catered to SNP variation, and WGS in particular has greatly expanded our view of genome wide variability within and between species.The increasing accessibility and coverage of WGS data has also enabled the direct identification of larger genetic variants, known as structural variants (SVs) (Alkan et al., 2011), enabling a deeper understanding of their role in ecology and evolution (Mérot, Oomen, et al., 2020). The growing breadth of population level SV studies has quickly revealed both the ubiquity and magnitude of SVs’ contributions towards intraspecific genomic diversity across species. A higher proportion of the genome is covered by SVs than SNPs (e.g. 5x (Mérot et al., 2023), 3x (Catanach et al., 2019), 8x (Hämälä et al., 2021)). SVs can also strongly affect fitness traits (e.g., 24% increased heritability estimates (Zhou et al., 2022)). Continued research into SVs will undoubtedly reveal more about the fundamental mechanisms of evolution across the tree of life. The inclusion of SVs into research fields that have previously focused heavily on SNPs will aid with the interpretation of complex genomic patterns and processes (Dallaire et al., 2023; Oomen et al., 2020) and a more complete picture of intraspecific and interspecific genetic variation.Because of this growing appreciation of SVs, researchers are increasingly interested in reanalysing existing WGS datasets or obtaining new data to examine SVs in their species of interest. This can be a daunting task because of the genetic resources required, as well as the technical and biological complexity of SV data analysis and interpretation. Furthermore, SVs are an extremely diverse category of variants, and analysing SVs like SNPs - as a single group - limits insights from their diverse subtypes and complex roles in genetic diversity.In this review, we aim to answer practical questions for those new to the study of SVs, guide study design and analytical best practices for those seeking to analyse population level SVs within eco-evolutionary studies, and suggest future avenues of inquiry. First, we summarise the differences between SVs and SNPs, as well as between diverse types of SVs, and discuss how these specific properties may interact with eco-evolutionary processes. We then focus on the practical side of SV analysis, providing a framework for identifying and analysing SVs from WGS data. Accompanying the global movement providing high quality reference genomes (e.g. Ebenezer et al., 2022; Lewin et al., 2022; Mc Cartney et al., 2024), this review aims to reduce barriers to analyzing the full spectrum of genomic variation by incorporating SVs in eco-evolutionary studies.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,019 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,003 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,004 | 0,007 |
| Science ouverte | 0,004 | 0,003 |
| Intégrité de la recherche | 0,005 | 0,011 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,098 | 0,080 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».