Automated seafloor massive sulfide detection through integrated image segmentation and geophysical data analysis: Revisiting the TAG hydrothermal field
Bibliographic record
Abstract
Accessible seafloor minerals located near mid-ocean ridges are noticed to mitigate projected metal demands of the net-zero energy transition, promoting growing research interest in quantifying global distributions of seafloor massive sulfides (SMS). Mineral potentials are commonly estimated using geophysical and geological data that lastly rely on additional confirmation studies using sparsely available, locally limited, seafloor imagery, grab samples, and coring data. This raises the challenge of linking in-situ confirmation data to geophysical data acquired at disparate spatial scales to obtain quantitative mineral predictions. Although multivariate datasets for marine mineral research are incessantly acquired, robust, integrative data analysis requires cumbersome workflows and experienced interpreters. Here, we introduce an automated two-step machine learning approach that integrates automated mound detection with geophysical data to merge mineral predictors into distinct classes and reassess marine mineral potentials for distinct regions. The automated workflow employs a U-Net convolutional neural network to identify mound-like structures in bathymetry data and distinguishes different mound classes through classification of mound architectures and magnetic signatures. Finally, controlled source electromagnetic data is utilized to reassess predictions of potential SMS volumes. Our study focuses on the Trans-Atlantic Geotraverse (TAG) area, which is amid the most explored SMS area worldwide and includes 15 known SMS sites. The automated workflow classifies 14 of the 15 known mounds as exploration targets of either high- or medium-priority. This reduces the exploration area to less than 7% of the original survey area from 49 km2 to 3.1 km2.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".