Development of a high-throughput single nucleotide polymorphism panel for genetic stock identification of Striped Bass
Bibliographic record
Abstract
ABSTRACT Objective The Striped Bass Morone saxatilis is an anadromous fish that has experienced population declines and recoveries throughout its range from the 1960s to the present day. While many United States fisheries have reopened since the most recent declines, ongoing monitoring is required to ensure that numbers remain stable in the long term. Central to this is the need to determine the extent to which major spawning locations contribute to coastal stocks. While next-generation genomic sequencing can discriminate among closely related breeding aggregations within the Striped Bass native range, the cost per sample of next-generation sequencing methods currently used is too high to be applied to large-scale and long-term projects moving forward. Methods We developed, optimized, and evaluated a GT-Seq panel—a small panel of highly informative single nucleotide polymorphisms—capable of assigning large numbers of Striped Bass to six genetically distinct regions across the Striped Bass spawning range at much lower cost than the previous next-generation sequencing methods. Results The final panel of 233 loci was able to assign 95% of reference individuals to region of origin and had <5% error across all simulations when estimating the mixing proportion of any stock. Conclusions This panel is being used in ongoing characterizations of Striped Bass along the Massachusetts coast and the eastern coast of Nova Scotia. For researchers tracking the migration of Striped Bass throughout their native range, this panel provides a low-cost method of genetically characterizing stocks at specific locations and times that can be easily modified or updated as additional, new genetic data become available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".