Integrating design- and model-based inference to estimate length and age composition in North Pacific longline catches
Bibliographic record
Abstract
Age and size structure are attributes of fishery stocks important for predicting future productivity. As such, estimating age and length composition of catches has long been an important fisheries management activity. Many observer programs sample catches to obtain length measurements and otoliths (or other structures for ageing) from targeted species. In North Pacific groundfish fisheries, observers collect these data through a stratified multiphase sampling design. Sampling variance and covariance estimates for catch- or proportions-at-length or -age that reflect the randomization inherent in the sampling design provide important measures of uncertainty that correspond to measurement error components in length- or age-structured stock assessment models. We compare sampling variances and covariances of Pacific cod (Gadus macrocephalus) proportions-at-length and sablefish (Anoplopoma fimbria) proportions-at-age with those provided by the overdispersed multinomial model sometimes used in these assessment models. For example, the sampling variance estimates for 2002 Pacific cod proportion-at-length estimates in the Bering Sea Aleutian Islands are at most 13% of the variances provided by multinomial and square-root sample size assumptions. Furthermore, some proportion estimates are positively correlated, whereas only negative correlation occurs with the multinomial distribution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.052 | 0.149 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".