Design and validation of a high-density single nucleotide polymorphism array for the Eastern oyster (<i>Crassostrea virginica</i>)
Bibliographic record
Abstract
Dense single nucleotide polymorphism (SNP) arrays are essential tools for rapid high-throughput genotyping for many genetic analyses, including genomic selection and high-resolution population genomic assessments. We present a high-density (200 K) SNP array developed for the Eastern oyster (Crassostrea virginica), which is a species of significant aquaculture production and restoration efforts throughout its native range. SNP discovery was performed using low-coverage whole-genome sequencing of 435 F1 oysters from families from 11 founder populations in New Brunswick, Canada. An Affymetrix Axiom Custom array was created with 219,447 SNPs meeting stringent selection criteria and validated by genotyping more than 4,000 oysters across 2 generations. In total, 144,570 SNPs had a call rate >90%, most of which (96%) were polymorphic and were distributed across the Eastern oyster reference genome, with similar levels of genetic diversity observed in both generations. Linkage disequilibrium was low (maximum r2 ∼0.32) and decayed moderately with increasing distance between SNP pairs. Taking advantage of our intergenerational data set, we quantified Mendelian inheritance errors to validate SNP selection. Although most of SNPs exhibited low Mendelian inheritance error rates overall, with 72% of called SNPs having an error rate of <1%, many loci had elevated Mendelian inheritance error rates, potentially indicating the presence of null alleles. This SNP panel provides a necessary tool to enable routine application of genomic approaches, including genomic selection, in C. virginica selective breeding programs. As demand for production increases, this resource will be essential for accelerating production and sustaining the Canadian oyster aquaculture industry.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".