Same Salmon Shared Semantics; Cross-community Salmon Data Standards for Data Integration and Decision Support
Notice bibliographique
Résumé
Salmon decisions stall on semantics, not on science. Take “wild salmon”: locally it can mean natural-origin fish, fish spawning naturally this year (including hatchery-origin spawners), or simply adipose-intact fish—definitions that change counts and benchmarks and challenge regional analyses. This fragmentation slows management, obscures accountability, and undermines confidence in otherwise excellent science. What’s needed is a shared vocabulary and an agreed-upon map of salmon terms—clear definitions and relationships that connect local labels to common meanings so people and software interpret data the same: a shared dictionary and rulebook for salmon data, an ontology. The DFO Salmon Ontology provides that map of how terms relate, and the controlled vocabularies that underpin it supply precise definitions—showing where terms differ, how they align, and where they should converge. Together, they standardize key terms across programs and regions. Teams can map local terms once, keep source systems unchanged yet aligned regionally, and link inputs to methods, benchmarks, and policy thresholds for Fisheries Science Reports, the Fish Stock Provisions, and the Wild Salmon Policy. Developed by the Fishery & Assessment Data Section in the Pacific Region Science Branch, this work builds on the International Year of the Salmon Data Mobilization initiative and collaborations with the National Center for Ecological Analysis and Synthesis (U.S.) and the global Research Data Alliance. It is open source and implements community standards from the W3C, OBO Foundry, and Darwin Core. By removing terminology friction, it prepares us for AI-assisted data integration and cross-discipline interoperability while immediately letting biologists spend less time cleaning data. Our goal is straightforward: to provide persistent, web-accessible definitions that help scientists and software combine data efficiently, support reproducible analyses, and strengthen confidence in salmon management.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,091 | 0,165 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,004 |
| Méta-épidémiologie (sens large) | 0,002 | 0,004 |
| Bibliométrie | 0,012 | 0,014 |
| Études des sciences et des technologies | 0,005 | 0,009 |
| Communication savante | 0,023 | 0,041 |
| Science ouverte | 0,010 | 0,028 |
| Intégrité de la recherche | 0,007 | 0,013 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,022 | 0,029 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».