Computational Oil-Slick Hub for Offshore Petroleum Studies
Bibliographic record
Abstract
The paper introduces the Oil-Slick Hub (OSH), a computational platform to facilitate the data visualization of a large database of petroleum signatures observed on the surface of the ocean with synthetic aperture radar (SAR) measurements. This Internet platform offers an information search and retrieval system of a database resulting from >20 years of scientific projects that interpreted ~15 thousand offshore mineral oil “slicks”: natural oil “seeps” versus operational oil “spills”. Such a Digital Mega-Collection Database consists of satellite images and oil-slick polygons identified in the Gulf of Mexico (GMex) and the Brazilian Continental Margin (BCM). A series of attributes describing the interpreted slicks are also included, along with technical reports and scientific papers. Two experiments illustrate the use of the OSH to facilitate the selection of data subsets from the mega collection (GMex variables and BCM samples), in which artificial intelligence techniques—machine learning (ML)—classify slicks into seeps or spills. The GMex variable dataset was analyzed with simple linear discriminant analyses (LDAs), and a three-fold accuracy performance pattern was observed: (i) the least accurate subset (~65%) solely used acquisition aspects (e.g., acquisition beam mode, date, and time, satellite name, etc.); (ii) the best results (>90%) were achieved with the inclusion of location attributes (i.e., latitude, longitude, and bathymetry); and (iii) moderate performances (~70%) were reached using only morphological information (e.g., area, perimeter, perimeter to area ratio, etc.). The BCM sample dataset was analyzed with six traditional ML methods, namely naive Bayes (NB), K-nearest neighbors (KNN), decision trees (DT), random forests (RF), support vector machines (SVM), and artificial neural networks (ANN), and the most effective algorithms per sample subsets were: (i) RF (86.8%) for Campos, Santos, and Ceará Basins; (ii) NB (87.2%) for Campos with Santos Basins; (iii) SVM (86.9%) for Campos with Ceará Basins; and (iv) SVM (87.8%) for only Campos Basin. The OSH can assist in different concerns (general public, social, economic, political, ecological, and scientific) related to petroleum exploration and production activities, serving as an important aid in discovering new offshore exploratory frontiers, avoiding legal penalties on oil-seep events, supporting oceanic monitoring systems, and providing valuable information to environmental studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".