sFDvent: A global trait database for deep‐sea hydrothermal‐vent fauna
Bibliographic record
Abstract
Abstract Motivation Traits are increasingly being used to quantify global biodiversity patterns, with trait databases growing in size and number, across diverse taxa. Despite growing interest in a trait‐based approach to the biodiversity of the deep sea, where the impacts of human activities (including seabed mining) accelerate, there is no single repository for species traits for deep‐sea chemosynthesis‐based ecosystems, including hydrothermal vents. Using an international, collaborative approach, we have compiled the first global‐scale trait database for deep‐sea hydrothermal‐vent fauna – sFDvent ( s Div‐funded trait database for the F unctional D iversity of vent s). We formed a funded working group to select traits appropriate to: (a) capture the performance of vent species and their influence on ecosystem processes, and (b) compare trait‐based diversity in different ecosystems. Forty contributors, representing expertise across most known hydrothermal‐vent systems and taxa, scored species traits using online collaborative tools and shared workspaces. Here, we characterise the sFDvent database, describe our approach, and evaluate its scope. Finally, we compare the sFDvent database to similar databases from shallow‐marine and terrestrial ecosystems to highlight how the sFDvent database can inform cross‐ecosystem comparisons. We also make the sFDvent database publicly available online by assigning a persistent, unique DOI. Main types of variable contained Six hundred and forty‐six vent species names, associated location information (33 regions), and scores for 13 traits (in categories: community structure, generalist/specialist, geographic distribution, habitat use, life history, mobility, species associations, symbiont, and trophic structure). Contributor IDs, certainty scores, and references are also provided. Spatial location and grain Global coverage (grain size: ocean basin), spanning eight ocean basins, including vents on 12 mid‐ocean ridges and 6 back‐arc spreading centres. Time period and grain sFDvent includes information on deep‐sea vent species, and associated taxonomic updates, since they were first discovered in 1977. Time is not recorded. The database will be updated every 5 years. Major taxa and level of measurement Deep‐sea hydrothermal‐vent fauna with species‐level identification present or in progress. Software format .csv and MS Excel (.xlsx).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".