Chum salmon baseline for the Western Alaska Salmon Stock Identification Program
Bibliographic record
Abstract
Uncertainty about the magnitude, frequency, location, and timing of the nonlocal harvest of sockeye and chum salmon was the impetus for the Western Alaska Salmon Stock Identification Program. The program was designed to use genetic data in mixed stock analysis to reduce this uncertainty. A baseline of allele frequencies in spawning populations is required for use in mixed-stock analysis to estimate the stock of origin of harvested fish. This report describes the methodology used to understand the population genetic structure among chum salmon populations and to build and test a baseline for use in mixed stock analysis of chum salmon. Of the 35,921 fish from 434 collections selected to be genotyped, the final baseline was composed of 32,817 fish from 402 collections representing 310 populations. Average population sample size was 106 fish. Reporting groups were determined through a combination of stakeholder needs and identifiability using genetic information, as measured using proof tests. The final reporting groups included Asia, Kotzebue Sound, Coastal Western Alaska, Upper Yukon River, Northern District (Alaska Peninsula), Northwest District (Alaska Peninsula), South Peninsula (Alaska Peninsula), Chignik/Kodiak, and East of Kodiak.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.023 | 0.016 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".