MétaCan
Menu
← Back to cohort
Record W7114796738 · doi:10.5281/zenodo.17883972

Same Salmon Shared Semantics; Cross-community Salmon Data Standards for Data Integration and Decision Support

2025· article· W7114796738 on OpenAlexaff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2025
Typearticle
Language
FieldEnvironmental Science
TopicFish Ecology and Management Studies
Canadian institutionsFisheries and Oceans Canada
Fundersnot available
KeywordsInteroperabilityTerminologySoftwareVocabularyWork (physics)Decision support systemData integrationBest practiceFragmentation (computing)

Abstract

fetched live from OpenAlex

Salmon decisions stall on semantics, not on science. Take “wild salmon”: locally it can mean natural-origin fish, fish spawning naturally this year (including hatchery-origin spawners), or simply adipose-intact fish—definitions that change counts and benchmarks and challenge regional analyses. This fragmentation slows management, obscures accountability, and undermines confidence in otherwise excellent science. What’s needed is a shared vocabulary and an agreed-upon map of salmon terms—clear definitions and relationships that connect local labels to common meanings so people and software interpret data the same: a shared dictionary and rulebook for salmon data, an ontology. The DFO Salmon Ontology provides that map of how terms relate, and the controlled vocabularies that underpin it supply precise definitions—showing where terms differ, how they align, and where they should converge. Together, they standardize key terms across programs and regions. Teams can map local terms once, keep source systems unchanged yet aligned regionally, and link inputs to methods, benchmarks, and policy thresholds for Fisheries Science Reports, the Fish Stock Provisions, and the Wild Salmon Policy. Developed by the Fishery & Assessment Data Section in the Pacific Region Science Branch, this work builds on the International Year of the Salmon Data Mobilization initiative and collaborations with the National Center for Ecological Analysis and Synthesis (U.S.) and the global Research Data Alliance. It is open source and implements community standards from the W3C, OBO Foundry, and Darwin Core. By removing terminology friction, it prepares us for AI-assisted data integration and cross-discipline interoperability while immediately letting biologists spend less time cleaning data. Our goal is straightforward: to provide persistent, web-accessible definitions that help scientists and software combine data efficiently, support reproducible analyses, and strengthen confidence in salmon management.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.091
metaresearch head score (Gemma)0.165
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.091
Threshold uncertainty score0.484

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0910.165
Meta-epidemiology (narrow)0.0020.004
Meta-epidemiology (broad)0.0020.004
Bibliometrics0.0120.014
Science and technology studies0.0050.009
Scholarly communication0.0230.041
Open science0.0100.028
Research integrity0.0070.013
Insufficient payload (model declined to judge)0.0220.029

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.070
GPT teacher head0.332
Teacher spread0.261 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)→Same topicFish Ecology and Management Studies→French-language works237,207→