Same Salmon Shared Semantics; Cross-community Salmon Data Standards for Data Integration and Decision Support
Bibliographic record
Abstract
Salmon decisions stall on semantics, not on science. Take “wild salmon”: locally it can mean natural-origin fish, fish spawning naturally this year (including hatchery-origin spawners), or simply adipose-intact fish—definitions that change counts and benchmarks and challenge regional analyses. This fragmentation slows management, obscures accountability, and undermines confidence in otherwise excellent science. What’s needed is a shared vocabulary and an agreed-upon map of salmon terms—clear definitions and relationships that connect local labels to common meanings so people and software interpret data the same: a shared dictionary and rulebook for salmon data, an ontology. The DFO Salmon Ontology provides that map of how terms relate, and the controlled vocabularies that underpin it supply precise definitions—showing where terms differ, how they align, and where they should converge. Together, they standardize key terms across programs and regions. Teams can map local terms once, keep source systems unchanged yet aligned regionally, and link inputs to methods, benchmarks, and policy thresholds for Fisheries Science Reports, the Fish Stock Provisions, and the Wild Salmon Policy. Developed by the Fishery & Assessment Data Section in the Pacific Region Science Branch, this work builds on the International Year of the Salmon Data Mobilization initiative and collaborations with the National Center for Ecological Analysis and Synthesis (U.S.) and the global Research Data Alliance. It is open source and implements community standards from the W3C, OBO Foundry, and Darwin Core. By removing terminology friction, it prepares us for AI-assisted data integration and cross-discipline interoperability while immediately letting biologists spend less time cleaning data. Our goal is straightforward: to provide persistent, web-accessible definitions that help scientists and software combine data efficiently, support reproducible analyses, and strengthen confidence in salmon management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.091 | 0.165 |
| Meta-epidemiology (narrow) | 0.002 | 0.004 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.012 | 0.014 |
| Science and technology studies | 0.005 | 0.009 |
| Scholarly communication | 0.023 | 0.041 |
| Open science | 0.010 | 0.028 |
| Research integrity | 0.007 | 0.013 |
| Insufficient payload (model declined to judge) | 0.022 | 0.029 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".