Ecological forecasts reveal limitations of common model selection methods: predicting changes in beaver colony densities
Bibliographic record
Abstract
Over the past two decades, there have been numerous calls to make ecology a more predictive science through direct empirical assessments of ecological models and predictions. While the widespread use of model selection using information criteria has pushed ecology toward placing a higher emphasis on prediction, few attempts have been made to validate the ability of information criteria to correctly identify the most parsimonious model with the greatest predictive accuracy. Here, we used an ecological forecasting framework to test the ability of information criteria to accurately predict the relative contribution of density dependence and density-independent factors (forage availability, harvest, weather, wolf [Canis lupus] density) on inter-annual fluctuations in beaver (Castor canadensis) colony densities. We modeled changes in colony densities using a discrete-time Gompertz model, and assessed the performance of four models using information criteria values: density-independent models with (1) and without (2) environmental covariates; and density-dependent models with (3) and without (4) environmental covariates. We then evaluated the forecasting accuracy of each model by withholding the final one-third of observations from each population and compared observed vs. predicted densities. Information criteria and our forecasting accuracy metrics both provided strong evidence of compensatory density dependence in the annual dynamics of beaver colony densities. However, despite strong within-sample performance by the most complex model (density-dependent with covariates) as determined using information criteria, hindcasts of colony densities revealed that the much simpler density-dependent model without covariates performed nearly as well predicting out-of-sample colony densities. The hindcast results indicated that the complex model over-fit our data, suggesting that parameters identified by information criteria as important predictor variables are only marginally valuable for predicting landscape-scale beaver colony dynamics. Our study demonstrates the importance of evaluating ecological models and predictions with long-term data and revealed how a known limitation of information criteria (over-fitting of complex models) can affect our interpretation of ecological dynamics. While incorporating knowledge of the factors that influence animal population dynamics can improve population forecasts, we suggest that comparing forecast performance metrics can likewise improve our knowledge of the factors driving population dynamics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".