Forecasting fish recruitment in age‐structured population models
Bibliographic record
Abstract
Abstract Recruitment in age‐structured stock assessment models can be forecasted using a variety of algorithms to provide advice on the anticipated consequences of different possible management actions. Selecting one method over another usually involves some subjectivity, yet can be consequential to the provision of advice. Extensive case‐specific testing is not always feasible. We evaluated the forecast skill in 3‐, 5‐ and 10‐year forecasts of 16 recruitment forecasting methods under various circumstances to provide a broad evaluation and general guidelines on the reliability of forecasts. We used 31 operating models based on existing stock assessment models applied to a diversity of stocks with empirical data, which we show to be generally representative of assessed stocks worldwide. Although no single best‐performing method could be identified, we found that time‐series methods were most likely to perform poorly. Both forecast skill across all methods and forecast sensitivity to the selected method were linked to the properties of the stock or assessment: age at maturity and recruitment autocorrelation in 3‐year forecasts and previous long‐term recruitment variability in 10‐year forecasts. In some situations, all forecasting methods resulted in systematic over‐ or underestimation of spawning stock biomass. The simulation approach employed here to assess forecast performance, rooted directly in the predictions of existing stock assessment models, can be a complementary tool to existing simulation approaches which generate alternative sets of population dynamics or observations and we discussed the advantages and limitations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".