In‐field genetic stock identification of overwintering coho salmon in the Gulf of Alaska: Evaluation of Nanopore sequencing for remote real‐time deployment
Bibliographic record
Abstract
Genetic stock identification (GSI) from genotyping-by-sequencing of single nucleotide polymorphism (SNP) loci has become the gold standard for stock of origin identification in Pacific salmon. The sequencing platforms currently applied require large batch sizes and multiday processing in specialized facilities to perform genotyping by the thousands. However, recent advances in third-generation single-molecule sequencing platforms, such as the Oxford Nanopore minION, provide base calling on portable, pocket-sized sequencers and promise real-time, in-field stock identification of variable batch sizes. Here we evaluate utility and comparability to established GSI platforms of at-sea stock identification of coho salmon (Oncorhynchus kisutch) using targeted SNP amplicon sequencing on the minION platform during a high-sea winter expedition to the Gulf of Alaska. As long read sequencers are not optimized for short amplicons, we concatenate amplicons to increase coverage and throughput. Nanopore sequencing at-sea yielded data sufficient for stock assignment for 50 out of 80 individuals. Nanopore-based SNP calls agreed with Ion Torrent-based genotypes in 83.25%, but assignment of individuals to stock of origin only agreed in 61.5% of individuals, highlighting inherent challenges of Nanopore sequencing, such as resolution of homopolymer tracts and indels. However, poor representation of assayed salmon in the queried baseline data set contributed to poor assignment confidence on both platforms. Future improvements will focus on lowering turnaround time and cost, increasing accuracy and throughput, as well as augmentation of the existing baselines. If successfully implemented, Nanopore sequencing will provide an alternative method to the large-scale laboratory approach by providing mobile small batch genotyping to diverse stakeholders.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".