Charlesworth <i>et al.</i> on Background Selection and Neutral Diversity
Bibliographic record
Abstract
A significant fraction of genetic diversity across the genome may be of little fitness consequence. But such neutral variation is profoundly informative about the evolutionary forces experienced by a population, and understanding what shapes this diversity across the genome is a key question in population genetics. In their landmark GENETICS article, Charlesworth et al. (1993) revealed an important effect structuring neutral genetic variation—a phenomenon called “background selection.” Their work forced the reinterpretation of influential empirical results and stimulated a major research program on how to distinguish background selection from other influences on neutral diversity. The question of what shapes neutral variation is closely linked to crucial questions about positive selection. How frequently does positive selection of adaptive variants occur, and how does this process influence different parts of the genome? In 1974, Maynard Smith and Haigh showed that positive selection will reduce neutral variation linked to the selected locus, an effect known as genetic “hitchhiking” (Maynard Smith and Haigh 1974). They predicted that, if positive selection is frequent enough, genetic diversity should be lower in regions of low recombination rates because, in these locations, beneficial mutations will be linked to a greater number of sites. This prediction set in motion empirical work that examined the distribution of neutral diversity in natural populations. Stephan and Langley (1989), Aguade et al. (1989), Begun and Aquadro (1992), and others showed that variation in the genomes of wild Drosophila species is indeed lower in regions of low recombination. This striking confirmation of Maynard Smith and Haigh’s predictions was widely taken as evidence for the action of frequent positive selection. But there was an alternative explanation. Charlesworth et al. described, for the first time, another effect that could produce lower neutral diversity in regions of low recombination: negative selection against deleterious mutations. Naming the phenomenon “background selection” they demonstrated that the effect could plausibly produce many of the observed patterns. Using classic results from the theory of mutation-selection balance, coupled with computer simulations that relax some of the simplifying assumptions, they showed that diversity in a nonrecombining region will be reduced as a function of the fraction of copies of the region that contain deleterious mutations. This is because copies that bear deleterious mutations are destined to be eliminated rapidly from the population, limiting possibilities for them to contribute to neutral diversity. In a large nonrecombining region, such as those surrounding centromeric regions, such a diversity loss can be substantial. While their analytical work assumed no recombination, their simulations explored the effects of partial recombination, and they were able to recover a correlation between recombination and diversity. Subsequent work by Hudson and Kaplan (1995), Nordborg et al. (1996), and others demonstrated how recombination rates can be incorporated into the equations predicting the amount of neutral diversity under background selection, quantifying how regions of very low recombination can be strongly affected by this process. This presented a new puzzle for population geneticists; negative and positive selection have very different implications for the evolutionary process, but seemed to have inconveniently similar effects on genetic diversity. Distinguishing between background selection and genetic hitchhiking became the focus of significant research and a vigorous debate that continues today. Unlike with positive selection, we have at least some direct experimental insights into the deleterious mutation rate, and the strength of selection against harmful mutations. These estimates mean we can attempt to predict and control for background selection in order to evaluate the evidence for an additional role for positive selection, as first shown by Charlesworth et al. (1993). As this study first demonstrated, both positive and negative selection are likely to jointly contribute to the structuring of neutral variation in Drosophila. Background selection is now widely acknowledged as a major force structuring genetic variation in many species, including humans, where it likely plays a major role in reducing variation near functional sites (e.g., Cai et al. 2009; McVicker et al. 2009; Lohmueller et al. 2011). It is also likely a major contributor to low genetic diversity on Y chromosomes, as well in asexual and selfing species (Glémin 2007; Agrawal and Hartfield 2016). Since background selection can increase the probability of fixing slightly deleterious mutations, this process can also contribute to the degeneration of the Y chromosome, and cause a decline in fitness of selfing and asexual lineages. The important advance made by Charlesworth et al. (1993) represents a remarkable case of feedback between theoretical and empirical population genetics, in which predictions from theory stimulated empirical tests, providing observations that motivated new theoretical work, forcing a rethink of the original observations and models, and prompting yet further advances in the continuing quest to understand the balance of evolutionary forces in natural populations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.005 | 0.006 |
| Insufficient payload (model declined to judge) | 0.063 | 0.020 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".