Remote Operating Vehicle Observation Data Contributions to the Ocean Biodiversity Information System
Bibliographic record
Abstract
In accordance with the United Nations Decade of Ocean Science for Sustainable Development, Ocean Networks Canada (ONC) is increasingly standardizing data in compliance with globally accepted schemas and concepts such as Essential Ocean Variables (EOV)s as directed by the Canadian Integrated Ocean Observing System (CIOOS). One such EOV is biological abundance data and ONC aims to provide this to end users as the widely accepted format - Darwin Core Archives (DwC-A). This will be done in collaboration with the UN’s Ocean Biodiversity Information System (OBIS). Although biological datasets are regularly contributed to OBIS, there are few provided by ROV visual observations. ONC hopes to provide a novel method for deep-sea ROV visual observation data as inspiration for similar dataset contributions in the future. By using the latest recommendations from peer reviewed sources, ONC will provide both transect-based and opportunistic biological observations as DwC-A files in event core format. These tables will provide various environmental sensor information as taken from the Remote Operating Vehicle (ROV) at the point of observation for enriching of observation data. Datasets will be organized in a hierarchical manner, with an individual dataset consisting of a single dive event. Relational information will be available in both the body of the DwC-A and the dataset’s metadata for linking back to parent expeditions as a dataset collection. ONC has developed conventions for the annotations of biological observations within ONC’s in-house solution for video annotation and viewing, SeaTubeV3. This modular web-based solution for annotations is based on taxonomies and button-sets. Button-sets in particular are a great way to map out a list of commonly seen organisms for quick and efficient event annotations. Coupled with a World Register for Marine Species (WoRMS) taxonomy and potential attribute pre-set assignment (individual counts, open nomenclature codes), SeaTubeV3 can provide a template for rich observation data collection with little need to cross-walk to different taxonomies downstream. Annotations generated from the dives will go through a verification process in which subject matter experts can vet whether a certain taxonomic identification is accurate or not. The annotations that meet a certain threshold will then be passed to ONC task machines where an OBIS-compliant DwCA package will be generated. These packages will be automated to populate on the OBIS web page and other repositories like CIOOS for anyone to use. By using these novel techniques, ONC hopes to help pave the way for similar deep-water biological observation datasets to be accessible through open formats like OBIS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".