Spatial modelling with Euclidean distance fields and machine learning
Bibliographic record
Abstract
Summary This study introduces a hybrid spatial modelling framework, which accounts for spatial non‐stationarity, spatial autocorrelation and environmental correlation. A set of geographic spatially autocorrelated Euclidean distance fields (EDF) was used to provide additional spatially relevant predictors to the environmental covariates commonly used for mapping. The approach was used in combination with machine‐learning methods, so we called the method Euclidean distance fields in machine‐learning (EDM). This method provides advantages over other prediction methods that integrate spatial dependence and state factor models, for example, regression kriging (RK) and geographically weighted regression (GWR). We used seven generic (EDFs) and several commonly used predictors with different regression algorithms in two digital soil mapping (DSM) case studies and compared the results to those achieved with ordinary kriging (OK), RK and GWR as well as the multiscale methods ConMap, ConStat and contextual spatial modelling (CSM). The algorithms tested in EDM were a linear model, bagged multivariate adaptive regression splines (MARS), radial basis function support vector machines (SVM), Cubist, random forest (RF) and a neural network (NN) ensemble. The study demonstrated that DSM with EDM provided results comparable to RK and to the contextual multiscale methods. Best results were obtained with Cubist, RF and bagged MARS. Because the tree‐based approaches produce discontinuous response surfaces, the resulting maps can show visible artefacts when only the EDFs are used as predictors (i.e. no additional environmental covariates). Artefacts were not obvious for SVM and NN and to a lesser extent bagged MARS. An advantage of EDM is that it accounts for spatial non‐stationarity and spatial autocorrelation when using a small set of additional predictors. The EDM is a new method that provides a practical alternative to more conventional spatial modelling and thus it enhances the DSM toolbox. Highlights We present a hybrid mapping approach that accounts for spatial dependence and environmental correlation. The approach is based on a set of generic Euclidean distance fields (EDF). Our Euclidean distance fields in machine learning (EDM) can model non‐stationarity and spatial autocorrelation. The EDM approach eliminates the need for kriging of residuals and produces accurate digital soil maps.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".