Quantifying habitat associations in marine fisheries: a generalization of the KolmogorovSmirnov statistic using commercial logbook records linked to archived environmental data
Bibliographic record
Abstract
Understanding specieshabitat associations is critical for designing marine reserves, defining essential fish habitat, and predicting the impacts of climate change on fisheries. For many species, however, there is a paucity of fisheries-independent data that simultaneously track abundance and environmental variables, as is the case for widow rockfish (Sebastes entomelas), a commercially important fishery off the west coast of the United States. In this paper, I generalize a previous approach to identifying habitat associations so that fisheries-dependent data can be used. In analyzing Oregon commercial logbook records and archived environmental data from the National Oceanographic Data Center, I found three environmental variables (bottom depth, vertical depth of fish in the water column, and temperature) to be statistically adequate. Using a generalized KolmogorovSmirnov test statistic, I compared an empirically derived cumulative distribution function (CDF) of the habitat sampled to a CDF weighted by widow rockfish catch. Results suggest that the significant habitat association for widow rockfish includes bottom depths between 136 and 298 m, vertical depths between 101 and 197 m, and temperatures between 7.1 and 8.1°C. This novel use of commercial logbook data, which links disparate data sources and explicitly accounts for unequal spatial sampling, is a methodological advance that also provides initial insights into widow rockfish habitat preferences.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.117 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".