Evaluating the performance of Gulf of Alaska walleye pollock (Theragra chalcogramma) recruitment forecasting models using a Monte Carlo resampling strategy
Bibliographic record
Abstract
Multiple linear regressions (MLRs), generalized additive models (GAMs), and artificial neural networks (ANNs) were compared as methods to forecast recruitment of Gulf of Alaska walleye pollock ( Theragra chalcogramma ). Each model, based on a conceptual model, was applied to a 41-year time series of recruitment, spawner biomass, and environmental covariates. A subset of the available time series, an in-sample data set consisting of 35 of the 41 data points, was used to fit an environment-dependent recruitment model. Influential covariates were identified through statistical variable selection methods to build the best explanatory recruitment model. An out-of-sample set of six data points was retained for model validation. We tested each model’s ability to forecast recruitment by applying them to an out-of-sample data set. For a more robust evaluation of forecast accuracy, models were tested with Monte Carlo resampling trials. The ANNs outperformed the other techniques during the model fitting process. For forecasting, the ANNs were not statistically different from MLRs or GAMs. The results indicated that more complex models tend to be more susceptible to an overparameterization problem. The procedures described in this study show promise for building and testing recruitment forecasting models for other fish species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".