Optimal Cross‐Validation Strategies for Selection of Spatial Interpolation Models for the Canadian Forest Fire Weather Index System
Bibliographic record
Abstract
Abstract The Canadian forest fire weather index (FWI) system requires spatially continuous, gridded weather data for temperature, relative humidity, wind speed, and precipitation. Reliable estimates of the Canadian FWI system components are needed to ensure the safety of communities, resources, and ecosystems. The quality of the interpolated input weather variables are typically evaluated using error estimates from cross‐validation. These error estimates are used for selecting between spatial interpolation methods for generating the continuous weather surfaces. Leave‐one‐out cross‐validation (LOOCV) is the most commonly used method, but it is biased in spatially clustered weather station networks. Accurate error estimation is important for selecting the optimal interpolation method and evaluating how well an interpolated surface represents true patterns in a weather variable. Other cross‐validation methods may better account for bias relating to clustered weather station networks. We present a comparison of cross‐validation methods for evaluating spatial interpolation models of weather variables for generating the inputs to the Canadian FWI system with the objective of determining whether they identify the same spatial interpolation model as having the lowest error. We found that LOOCV, shuffle‐split, stratified shuffle‐split, and a modified buffered leave‐one‐out procedure generally identified the same spatial interpolation models as having the lowest error. Spatial k‐fold favored spatial interpolation models with extrapolation ability. Our findings indicate that the most computationally efficient cross‐validation approach can be used for automatically selecting spatial interpolation models for weather surface generation, which will improve the quality of historical daily FWI maps.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.036 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".