EUNIS habitat classification and associated confidence shapefiles for Atlantic regions investigated by the EU H2020 project iAtlantic (Version 2)
Bibliographic record
Abstract
This dataset includes 11 regional EUNIS-classified habitat maps (100-1000 km) and associated confidence maps that were created as a project milestone (Nr. 12) of the EU H2020 project 'iAtlantic'. The 12 iAtlantic regions encompass 1. Subpolar Mid-Atlantic Ridge, off Iceland MFRI, 2. Rockall Trough to PAP, 3. Central mid-Atlantic Ridge, 4. NW Atlantic, Gully Canyon, 5. Sargasso Sea, 6. Eastern Tropical North Atlantic, Cape Verde, 7. Equatorial Atlantic, Romanche Fracture Zone, 8. Slope & margin off Angola & Congo Lobe, 9. Benguela Current, Walvis Ridge to South Africa, 10. Brazil margin & Santos and Campos Basin, 11. Vitória-Trindade Seamount Chain and 12. Malvinas Current. For each of the regions 2-12, a shapefile of polygons classified according to the 2022 EUNIS classification level 3 and a second shapefile of the same polygons attributed with their confidence level according to the MESH Accuracy & Confidence Working approach was created. EUNIS classifications combined biozone and substrate data. Biozones were assigned from bathymetry. Where MBES was not available, GEBCO bathymetry was used. Substrate data were extracted from pre-existing geological/substrate mapping efforts and converted to EUNIS classifications via cross walks or, where substrate data were limited, substrate layers were modelled using Random Forest. The EUNIS habitat map for Region 4 was based on the pre-existing surficial geology compilation of the Scotian Shelf bioregion compiled by the Geological Survey of Canada. The EUNIS habitat map for Region 9 was based on the pre-existing South African habitat map that uses a modified IUCN hierarchical classification system. No additional information to that used in the EUSeaMap was available for Region 1. Therefore, shapefiles were not created for Region 1.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.041 | 0.026 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".