Becoming pure: identifying generational classes of admixed individuals within lesser and greater scaup populations
Bibliographic record
Abstract
Estimating the frequency of hybridization is important to understand its evolutionary consequences and its effects on conservation efforts. In this study, we examined the extent of hybridization in two sister species of ducks that hybridize. We used mitochondrial control region sequences and 3589 double-digest restriction-associated DNA sequences (ddRADseq) to identify admixture between wild lesser scaup (Aythya affinis) and greater scaup (A. marila). Among 111 individuals, we found one introgressed mitochondrial DNA haplotype in lesser scaup and four in greater scaup. Likewise, based on the site-frequency spectrum from autosomal DNA, gene flow was asymmetrical, with higher rates from lesser into greater scaup. However, using ddRADseq nuclear DNA, all individuals were assigned to their respective species with >0.95 posterior assignment probability. To examine the power for detecting admixture, we simulated a breeding experiment in which empirical data were used to create F1 hybrids and nine generations (F2-F10) of backcrossing. F1 hybrids and F2, F3 and most F4 backcrosses were clearly distinguishable from pure individuals, but evidence of admixed histories was effectively lost after the fourth generation. Thus, we conclude that low interspecific assignment probabilities (0.011-0.043) for two lesser and nineteen greater scaup were consistent with admixed histories beyond the F3 generation. These results indicate that the propensity of these species to hybridize in the wild is low and largely asymmetric. When applied to species-specific cases, our approach offers powerful utility for examining concerns of hybridization in conservation efforts, especially for determining the generational time until admixed histories are effectively lost through backcrossing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".