Delineation of significant benthic areas in eastern Canada using kernel density analysis and species distribution models
Bibliographic record
Abstract
The Canadian Policy for Managing the Impact of Fishing on Sensitive Benthic Areas developed by the Department of Fisheries and Oceans Canada (DFO) in 2009 defines Significant Benthic Areas in DFO's Ecological Risk Assessment Framework as "significant areas of cold-water corals and sponge dominated communities". Kernel density estimation (KDE) was applied to research vessel trawl survey data to create modelled biomass surfaces for corals and sponges. From these, an aerial expansion method was applied to identify significant concentrations of these taxa across eastern Canada. The borders of the areas so identified were refined using species distribution models that predict species presence-absence and/or biomass, both incorporating environmental data. We present such predictive models produced using a random forest (RF) machine-learning technique. A suite of between 54 and 78 environmental predictor variables from different data sources were used. Occurrence models performed well in general with cross-validated AUC (Area Under the Receiver Operating Characteristic Curve) values over 0.8 in most of the cases. Biomass models provided diverse results depending of the taxa and region studied. The biomass models were compared with Generalized Additive Models (GAM), which produced comparable results to random forest, although the fewer assumptions required for RF made this method more convenient. These results have been used to identify significant concentrations of corals and sponges in eastern Canada, an essential first step in the identification of Sensitive Benthic Areas to ensure Canadian fisheries are conducted in a manner that supports marine conservation and sustainable resource use within and outside Canada's 200 nautical mile exclusive economic zone.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".