Evaluating SNP ascertainment bias and its impact on population assignment in Atlantic cod, <i>Gadus morhua</i>
Bibliographic record
Abstract
The increasing use of single nucleotide polymorphisms (SNPs) in studies of nonmodel organisms accentuates the need to evaluate the influence of ascertainment bias on accurate ecological or evolutionary inference. Using a panel of 1641 expressed sequence tag-derived SNPs developed for northwest Atlantic cod (Gadus morhua), we examined the influence of ascertainment bias and its potential impact on assignment of individuals to populations ranging widely in origin. We hypothesized that reductions in assignment success would be associated with lower diversity in geographical regions outside the location of ascertainment. Individuals were genotyped from 13 locations spanning much of the contemporary range of Atlantic cod. Diversity, measured as average sample heterozygosity and number of polymorphic loci, declined (c. 30%) from the western (H(e) = 0.36) to eastern (H(e) = 0.25) Atlantic, consistent with a signal of ascertainment bias. Assignment success was examined separately for pools of loci representing differing degrees of reductions in diversity. SNPs displaying the largest declines in diversity produced the most accurate assignment in the ascertainment region (c. 83%) and the lowest levels of correct assignment outside the ascertainment region (c. 31%). Interestingly, several isolated locations showed no effect of assignment bias and consistently displayed 100% correct assignment. Contrary to expectations, estimates of accurate assignment range-wide using all loci displayed remarkable similarity despite reductions in diversity. Our results support the use of large SNP panels in assignment studies of high geneflow marine species. However, our evidence of significant reductions in assignment success using some pools of loci suggests that ascertainment bias may influence assignment results and should be evaluated in large-scale assignment studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".