Comparing, evaluating and combining statistical species distribution models and <scp>CLIMEX</scp> to forecast the distributions of emerging crop pests
Bibliographic record
Abstract
BACKGROUND: Forecasting the spread of emerging pests is widely requested by pest management agencies in order to prioritise and target efforts. Two widely used approaches are statistical Species Distribution Models (SDMs) and CLIMEX, which uses ecophysiological parameters. Each have strengths and weaknesses. SDMs can incorporate almost any environmental condition and their accuracy can be formally evaluated to inform managers. However, accuracy is affected by data availability and can be limited for emerging pests, and SDMs usually predict year-round distributions, not seasonal outbreaks. CLIMEX can formally incorporate expert ecophysiological knowledge and predicts seasonal outbreaks. However, the methods for formal evaluation are limited and rarely applied. We argue that both approaches can be informative and complementary, but we need tools to integrate and evaluate their accuracy. Here we develop such an approach, and test it by forecasting the potential global range of the tomato pest Tuta absoluta. RESULTS: The accuracy of previously developed CLIMEX and new statistical SDMs were comparable on average, but the best statistical SDM techniques and environmental data substantially outperformed CLIMEX. The ensembled approach changes expectations of T. absoluta's spread. The pest's environmental tolerances and potential range in Africa, the Arabian Peninsula, Central Asia and Australia will be larger than previous estimates. CONCLUSION: We recommend that CLIMEX be considered one of a suite of SDM techniques and thus evaluated formally. CLIMEX and statistical SDMs should be compared and ensembled if possible. We provide code that can be used to do so when employing the biomod suite of SDM techniques. © 2021 The Authors. Pest Management Science published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.021 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".