Vulnerable Marine Ecosystem Indicator Taxa recorded by submarine as evidence of the presence of Vulnerable Marine Ecosystems, Antarctic Peninsula - data
Bibliographic record
Abstract
“Vulnerable Marine Ecosystem Indicator Taxa recorded by submarine as evidence of the presence of Vulnerable Marine Ecosystems, Antarctic Peninsula - data” is a sampling event type dataset published by AntOBIS. This resource supplements the publications listed in the bibliographic citation section. This dataset contains records of taxonomic groups that are considered VME-IT by the Commission for the Conservation of Antarctic Marine Resources (CCAMLR) and their relative percent abundances compared to non-VME-IT and bare substrate based on video footage captured by submarine deployed by the MY Arctic Sunrise during their Antarctica expeditions. The first took place in 2018 and focused within the Gerlache Strait and along the western Antarctic Peninsula and the Antarctic Sound in January 2018. Dives were conducted beginning 19th to 27th January 2018. A second expedition took place in 2022 and focused along the eastern Antarctic Peninsula in the Vegas Basin and the Erebus and Terror Gulf. Dives were conducted from 26th February to 6th March 2022. The data are published as a standardized Darwin Core Archive and includes locality, coordinates, event date, depth, sampling protocol, sampling effort, occurrence status, vernacular name, scientific name and taxa classification. Multimedia associated with this resource is available at Zenodo: https://doi.org/10.5281/zenodo.6760063 This dataset is published by SCAR-AntOBIS under the license CC-BY 4.0. Please follow the guidelines from the SCAR and IPY Data Policies (https://www.scar.org/excom-meetings/xxxi-scar-delegates-2010-buenos-aires-argentina/4563-scar-xxxi-ip04b-scar-data-policy/file/) when using the data. If you have any questions regarding this dataset, please contact us via the contact information provided in the metadata or via data-biodiversity-aq@naturalsciences.be. Issues with dataset can be reported at https://github.com/biodiversity-aq/data-publication/ This dataset is part of the Southern Benthics VME project in conjunction with Greenpeace International which conducted the expedition.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.058 | 0.035 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".