MétaCan
Menu
← Retour à la cohorte
Enregistrement W4410852767 · doi:10.22541/au.174853973.36642913/v1

A Beginner’s Guide to Structural Variants in Eco-Evolutionary Population Genomics: Everything You Wanted to Know

2025· preprint· en· W4410852767 sur OpenAlexafffund
Katarina C. Stuart, Rebekah A. Oomen, Maren Wellenreuther, Jana Wold, Anna Tigano, David L. Field, Claire Mérot

Notice bibliographique

Revuenon disponible
Typepreprint
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueEvolution and Genetic Dynamics
Établissements canadiensFisheries and Oceans CanadaUniversity of New Brunswick
Organismes subventionnairesNatural Sciences and Engineering Research Council of CanadaVetenskapsrådetMinistry of Business, Innovation and EmploymentFisheries and Oceans CanadaNew Brunswick Innovation Foundation
Mots-clésPopulation genomicsGenomicsNeed to knowPopulationEvolutionary biologyBiologyData scienceComputer scienceGeneticsSociologyGenomeGeneDemography

Résumé

récupéré en direct d'OpenAlex

Whole genome sequencing (WGS) has greatly expanded researchers' ability to study structural variants (SVs), i.e. the variation in the presence, number, orientation, or position of a DNA sequence. This has paved the way to study the eco-evolutionary dynamics of SVs across the tree of life and within a population genomics framework. In this review, we provide the necessary fundamentals to help researchers generate and analyze population-level SV data. We discuss the unique properties of different SV groups, and how these fundamental differences interact with important biological and evolutionary processes using both empirical results and theory. This includes discussion of unresolved issues around SVs, such as technical difficulties in identification, accounting for diversity, and evaluating functional effects. We explicitly integrate into this discussion transposable elements, which are an important component of SVs often identified in population-level variant data. Finally, we focus on the practical side of SV analysis, offering a framework for SV identification and data analysis. In particular, we examine the heterogeneous nature of SV properties (type, length, sequence identity) that should be considered when studying them in ecology and evolution. This review aims to provide resources and guidelines to help researchers, navigate the complexities of a relatively new field of eco-evolutionary genomics research.Keywords: inversions, chromosomal rearrangements, copy number variants, transposable elements, distribution of fitness effects, rapid adaptation 1. IntroductionCharacterizing genomic variation is fundamental to address a wide array of ecological and evolutionary questions. The advancement of DNA sequencing methods over time has enabled the discovery of new aspects of genetic diversity at every step. In particular, ecological and evolutionary genomics have flourished with the increased availability of high quality reference genomes (Formenti et al., 2022), and attainability of whole genome sequencing (WGS) (Fuentes-Pardo & Ruzzante, 2017). Resequencing entire genomes, rather than a small portion through reduced-representation approaches, provides a rich source of information and has led to a proliferation of methods to investigate evolutionary and demographic processes. We can now identify signatures of balancing selection in the genome (Stern & Lee, 2020), reconstruct demographic history in the near and distant past with unprecedented resolution (Nadachowska-Brzyska et al., 2022; Santiago et al., 2020), and characterize the roles of genome structure and recombination in the levels and distribution of genomic variation across the genome (Akopyan et al., 2025; Tigano et al., 2021). Many genomic analysis methods are currently catered to SNP variation, and WGS in particular has greatly expanded our view of genome wide variation within and between species.The increasing accessibility and coverage of WGS data has also enabled the direct identification of larger genetic variants, known as structural variants (SVs) (Alkan et al., 2011), enabling a deeper understanding of their role in ecology and evolution (Mérot, Oomen, et al., 2020). The growing breadth of population-level SV studies has quickly revealed both the ubiquity and magnitude of SVs’ contributions towards genomic diversity. Several population genomics studies have shown that SVs generally cover more base pairs than sequence variation (i.e., Single Nucleotide Polymorphisms, SNPs) by a factor of 3x to 8x across investigated species (Catanach et al., 2019, Mérot et al., 2023, Hämälä et al., 2021, Tigano et al. 2020)). SVs can also strongly affect fitness and phenotypes (Wellenreuther et al., 2025). For example, the estimated heritability of agronomically-important traits in tomatoes increased by 24% when SVs were also considered compared to analyses based on SNPs only (Zhou et al., 2022). The inclusion of SVs into research fields that have previously focused heavily on SNPs will aid with the interpretation of complex genomic patterns and processes (e.g., biogeography, Dallaire et al., 2023; eco-evolutionary dynamics, Oomen et al., 2020) and provide researchers with a more complete picture of intraspecific and interspecific genetic variation.Because of this growing appreciation of SVs, researchers are increasingly interested in reanalysing existing WGS datasets or obtaining new data to examine SVs in their species of interest. This can be a daunting task because of the genetic resources required, as well as the technical and biological complexity of SV data analysis and interpretation. Furthermore, SVs are an extremely diverse category of variants, and analysing SVs like SNPs - as a single group - limits insights from their diverse subtypes and complex roles in genetic diversity.In this review, we aim to answer practical questions for those new to the study of SVs, guide study design and analytical best practices for those seeking to analyse population level SVs within eco-evolutionary studies, and suggest future avenues of inquiry. First, we summarise the differences between SVs and SNPs, as well as between diverse types of SVs, and discuss how these specific properties may interact with eco-evolutionary processes. We then focus on the practical side of SV analysis, providing a framework for identifying and analysing SVs from WGS data. Accompanying the global movement providing high quality reference genomes (e.g. Ebenezer et al., 2022; Lewin et al., 2022; Mc Cartney et al., 2024), this review aims to reduce barriers to analysing the full spectrum of genomic variation by incorporating SVs in eco-evolutionary studies.2. How do we define and classify SVs, including TEs?‘Structural Variant’ is a broad term that encompasses all variation in the DNA sequence other than single nucleotide variants (SNV, which include single nucleotide polymorphisms [SNPs]). SVs are generally defined as variation in the presence, absence, number, orientation, or position of a DNA sequence. Some studies classify a variant as structural if it exceeds a minimum length threshold (typically 50 bp), which is usually an arbitrary cutoff inherited from SV detection software (Mérot, Oomen, et al., 2020). At the extreme, SVs can be whole chromosomes and genome duplications (Scherer et al., 2007). Earlier methods for identifying small structural variants, such as insertion–deletions (indels), were limited by short-read sequencing and ignored variants longer than this threshold. Modern SV detection tools now often target this intermediate length range. In reality, however, such variants occur along a continuous length spectrum, extending from just a few base pairs to many megabases (Mérot, Oomen, et al., 2020, Wellenreuther et al. 2025). Any length threshold is thus inherently arbitrary and constrains the detection and interpretation of evolutionarily relevant SVs (Recuerda & Campagna, 2024). In practice, the apparent length range of SVs in any given study is determined not by biology but by the technical limits of variant-calling algorithms and the characteristics of the sequencing data. Longer reads generally enable the detection of larger and more complex variants (Mahmoud et al., 2019). Consequently, good scientific practice requires authors to clearly specify the length range of SVs targeted in their analyses, ideally grounded in the empirically defined detection limits established by benchmarking studies of the chosen SV caller(s) (e.g., Helal et al., 2024). This is particularly important when the operational definition of ‘SV’ in a study depends on those methodological constraints. For example, studies based on assembly alignment will perform well at characterising variants spanning hundreds of kilobases, whereas short-read based approaches capture those below a few kilobases (He et al., 2025). Longread (LR) platforms such as PacBio HiFi or Oxford Nanopore are theoretically able to recover variants of any length, but it is worth keeping in mind that most of the tools have been designed and tested on simulated or curated databases with a majority of SVs between 50bp and a few kilobases.SVs are also defined by their sequence change relative to a reference genome. This usually includes deletions (DEL), insertions (INS), duplications , inversions (INV), fusions, and translocations (Alkan et al., 2011). While this categorization is meaningful from a bioinformatic perspective (the variant is classified by comparing with the reference genome), other intrinsic characteristics of SVs may be more relevant from a biological point of view, such as how these variants originate and evolve over time. SVs can originate by many mechanisms, including errors in meiotic recombination like incomplete crossover (e.g., due to age or toxins), improper DNA repair, or replication issues like template switching or slippage (Carvalho & Lupski, 2016; Currall et al., 2013). Some SVs (but not all), when their sequence is examined, will be identified as a repeat, for example microsatellites or transposable elements (TEs). All these major classifications of SVs can also be caused by the activity of TEs (Almojil et al., 2021), which are repetitive genetic elements that can originate from the genome itself or from external viruses and have the ability to move and replicate themselves within the genome (Bourque et al., 2018). When a TE replicates or relocates within the host genome, it creates structural genetic variation, thus making this TE an SV. While SVs are identified by their differences from a reference genome, TEs are identified by their recurrent sequence motifs, which are generally phylogenetically grouped into 'families' sharing similar sequences that diversify alongside their host genomes (Bourque et al., 2018). When TEs become fixed due to selection or drift within the population or species, they are no longer a polymorphic variant, so they should not e considered a SV despite showing TE-like sequences. Thus, not all SVs are TEs, and conversely, not all TEs are SVs. Whether TEs are SVs or not, they can promote SV formation by creating similar genomic regions that trigger non-allelic homologous recombination (Klein & O’Neill, 2018, Harringmeyer & Hoekstra, 2022; Meyer et al., 2024). In this review, we use 'SV' to refer collectively to structural variants of all types and lengths, including those of TE and non-TE origin, unless otherwise specified. Note that many studies focus on specific subsets of SVs, using different terms such as copy number variation (CNV), indels (insertions-deletions), presence-absence variation (PAV), chromosomal rearrangements (CRs, e.g. inversions, translocations, fusions, usually longer than 100s of kb), or microsatellites (which are CNVs, which are in turn indels). Similarly, TEs are referred to as transposons, jumping genes, mobile genetic elements (MGEs), mobile DNA, retrotransposons or DNA transposons. Overlooking this diverse terminology may lead to important research being missed.2.1 Why do we study SV diversity when we already have genome wide SNP data? One might question whether identifying SVs is necessary—can SNP sequence variation alone capture the patterns of genetic variation necessary for genomic analysis? Broadly, patterns of diversity and differentiation (e.g., population structure) across sequence (SNPs) and structural (SVs) variation often correlate (Tigano et al., 2024; Tigano & Russello, 2022), although differential patterns are sometimes observed (Dorant et al., 2020; Tigano et al., 2024). SVs can affect patterns of population structure if they harbour a concentration of highly differentiated SNPs that capture an axis of differentiation, for example local adaptation, different from the “neutral” patterns of population differentiation (Tepolt et al., 2022). Moreover, because large structural rearrangements often underlie ecotypic differentiation, they frequently harbour loci of large effect that contribute to local adaptation. Such variants are therefore important to consider in the context of species management and conservation (Wold et al., 2021, Schneller et al., 2025). Their relevance extends to population viability, as demonstrated in wolves (Canis lupus lupus), where structural genomic variation in the inbred Scandinavian population contributes to realized genetic load (the accumulation of deleterious variants) but can be mitigated by immigration (Smeds et al., 2024). Similarly, in Atlantic salmon (Salmo salar), population-specific structural variants have been associated with different local adaptation between ecotypes (Lecomte et al., 2024). Throughout the rest of this review we will discuss many more examples, across a wide variety of SV types and different taxa. Much of the theory and analytical approaches used in population genomics has been developed around SNPs, and expanding the genomic toolkit to include SVs is still in its infancy (e.g. Barton & Zeng, 2018). Due to their dense and genome-wide distribution, SNPs remain ideal for some applications such as linkage disequilibrium (LD) studies (e.g. inferring recombination landscapes, effective population size, etc) and for quantitative trait loci (QTL) mapping. However, because SVs are larger and encompass more base pair changes overall, limiting the population genomic inference to SNP-based analysis will misrepresent overall genetic diversity. SNPs located within SVs may not be independent markers meaning the patterns they capture may be overrepresented in population genetic data, and complex SVs are often not captured by SNPs at all. This limitation disproportionately affects the detection of large-effect variants in the genome, as larger, more complex SVs are more likely to have functional impacts on the organism (see Section 3.2). SVs can be the genomic basis of discrete morphotypes (Lamichhaney et al., 2016) and ecotypes (Li et al., 2024), and underlie many human diseases, likely contributing to the missing heritability issue when overlooked (i.e. where trait heritability estimates are much lower than expected when calculated using SNPs only) (Groza, Chen, et al., 2024). An increasing number of studies on commercially-relevant species demonstrate that SVs underlie traits of economic interest (Jayakodi et al., 2020; Leonard et al., 2024). Therefore, important sources of genomic variation will be missed in ecology, evolution, and in applied research if SVs are ignored.One may wonder to what extent the effect of SV can be predicted from neighbouring SNPs. This approach may be cost-efficient in some cases such as well-documented catalogues of variants (Blaj et al., 2022) or for diagnostic SNPs associated with large inversions (Fang & Edwards, 2024). However, in general, SNP calling pipelines often exclude SV signals inadvertently (or intentionally) by or SNPs with high or which SNP signals (e.g. Dallaire et al., 2023; et al., but SV SVs are often in regions of high recombination et al., et al., and may to selection than SNPs which will the patterns of linkage disequilibrium between SVs and SNPs et al., The of both biological and technical that whether SVs are to SNPs is likely highly on the SV properties (e.g., and genomic context et al., et al., are the properties of SVs that for population the population genomics of SVs requires their properties and how they can interact with eco-evolutionary processes genome-wide SVs from such as in on variants with large or SVs like deletions or large However, understanding how a study unique history with SV properties can guide the of evolutionarily relevant The and dynamics of SVs are diverse and formation of SVs, including TEs, have been at the within the context of and genome evolution et al., 2017). The diversity in genome across the tree of life the highly of SV and and their contributions to genome et al., The at which SVs is expected to be more than for SNPs, the wide diversity in SV types and & 2021). SVs are to at a lower than SNPs, on in terms of SNP formation (e.g. genome et al., SV formation can a number of However, some types of SVs may more frequently than SNPs (e.g. small microsatellites et al., or duplications & The of SVs that are TEs are known for their high activity levels & this across TE and across time et al., 2021). Many eco-evolutionary and population genomics for SNPs (e.g. et al., 2022), which is known to be a et al., 2023; et al., 2021). However, SVs are more evolutionarily than SNPs as they and evolve in a variety of different and any fixed and than when applied to sequence variation et al., 2022; TEs, for example, are well known to often in when by population or & to many new TE In such the data a of variants, which may be to population or which be an interpretation in this & 2019). when to from SV identifying their of (e.g., what of an is an analytical and such data alongside with SNP will likely evolutionary research insights into the between and demographic history SVs are likely to be due to their length that the larger an the the that it will a likely on fitness et al., et al., 2021). Thus, we may that the distribution of fitness (the range and of fitness of new from deleterious to to of SVs to have a lower of variants than SNPs The fitness will be the for SNPs and SVs (i.e. the variant is however, the of large SVs deleterious SVs may be more than deleterious SNPs 2019). the magnitude of effect of SVs most likely extends that of SNPs with many large polymorphisms being example of variation et al., The distribution of fitness the of formation genome and selection species can differences in the fitness of SNPs, and SVs may be more diverse in this et al., variants are more likely to have functional and thus effect generally with variant length et al., 2020), with SVs more to SNPs, et al., The fitness of an SV also on the of sequence with the or of genomic sequence likely to be of selection in different et al., et al., 2022). The of the variant is also for example, are likely to be compared to deletions Finally, the of an SV may change over time. All SVs, and inversions in particular, are to genetic to recombination and et al., 2021), which may the load they over et al., between length and fitness for SVs is also likely to interact with many biological processes in different compared to SNPs. variants may local adaptation in the of which can so SVs may a role than SNPs in such & 2011). SVs may also help local adaptation high as of variants can be by whereas the SV will variants within its length et al., 2024), and the SV itself may recombination (see Section for of differences in and fitness effects, the evolutionary dynamics of SVs may across population and species For example, within some duplications and translocations have been to with increasing they may be SV that whereas differences in inversions were more and highly et al., 2024; & Consequently, the different characteristics of SVs such as length and are important to alongside and demographic when to the fitness impacts of changes to the genome et al., 2020). However, due to the detection in many SV studies to (see Section a framework for population-level of variant lengths, and has to be SVs can population level dynamics in that SNPs SVs sequences when in the which may or recombination inversions are the most in this & other types of SVs, such as fusions, large or complex rearrangements, may have similar on this is by the or of or & & et al., 2019). SVs frequently in recombination et al., but can or recombination et al., 2017). with high recombination are more effective at deleterious variants et al., et al., however, selection within regions may the of SVs that to has on genome but population dynamics and the level of genetic variation 2013). the rest of the genome is by the of with or no between them and effective population (see et al., for SVs recombination may variants over often increasing deleterious et al., 2022; et al., et al., 2019). SVs can variants into particularly in inversions (Wellenreuther & 2018). recombination can also the of variants, such as from range to in expanding et al., the of of SVs can their overall fitness thus the of both within species, in adaptation, and across species, when barriers are still et al., 2024; et al., et al., 2025; et al., 2022). SVs can genome and in that SNPs SVs can the of DNA, for example by the of which to sequence such as target and their sequences et al., 2018). SVs may accessibility to thus the sequence itself (Bourque et al., 2018, et al., 2019). The role of SVs in accessibility can be using et al. 2013). most applications of this have focused on and studies, its use in evolutionary biology is For example, et al. that SVs for of regions with accessibility across Many of these SVs to transposable insertions from different TE species and were often located near of DNA for changes in is which between regions of the genome, and has been used to demonstrate changes in structure and which in turn the position of chromosomes the formation of SVs et al., 2021). In to small SVs may to genome through changes to or SVs that are TEs an important role due to their unique may in change because the host of may DNA structure or which may in turn in target regions (Klein & O’Neill, 2018). of the between TEs and their host genomes, of can lead to in of in by the of TEs, thus showing a role of TEs in the of barriers between et al., genome TE may also be through from selection for example, which of and genetic et al., et al., & O’Neill, 2018). TEs sequence of within the host genome (Bourque et al., sequences may also functional change by a in which TE DNA sequences are by the host genome in to creating new or previously et al., 2017). the evolutionary of TE variation can be by population genomics with data to identify TE with high et al., or identify from et al., such mechanisms, SVs, TEs, can both evolution and through the of functional variation the genome (e.g., et al., 2025).

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,015
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: aucune
Score de désaccord entre enseignants0,044
Score d'incertitude au seuil0,149

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0040,015
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0020,002
Bibliométrie0,0040,004
Études des sciences et des technologies0,0010,003
Communication savante0,0030,008
Science ouverte0,0040,003
Intégrité de la recherche0,0050,010
Charge utile insuffisante (le modèle a refusé de juger)0,0440,046

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,006
Tête enseignante GPT0,279
Écart entre enseignants0,273 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2025
Routes d'admission2
Résumé présentoui

Explorer davantage

Même sujetEvolution and Genetic Dynamics→Travaux en français237 207→