Advancing digital soil mapping with multi-year crop cover data: Impacts on model accuracy and soil interpretation
Bibliographic record
Abstract
Vegetation cover has a significant influence on soil properties and is commonly used as a covariate in digital soil mapping (DSM). Crop frequency (CrFr) covariates, representing the frequency with which a certain crop or class of crops are grown over multiple years, can be derived from multi-year vegetation data. Such data have the potential to provide promising insights into soil conditions and can enhance predictions of soil properties. Predictive modelling within a DSM framework can improve our understanding of the relationship between crop cover and different soil properties. This study had two main objectives: (1) to develop DSM models for six soil properties—bulk density (BD), organic carbon (OC), A horizon thickness (AT), total nitrogen (TN), pH, and cation exchange capacity (CEC)—both with and without CrFr covariates, and to compare their accuracy metrics; each soil property was modelled independently as a separate response variable; and (2) to investigate the relationships between covariates such as crop types, precipitation, and temperature and soil properties. The study was conducted in the Ottawa, Canada, region, an area with diverse crop cover. From 13 years of Annual Crop Inventory (ACI) raster data, five CrFr covariates were generated and added to other covariates commonly used in DSM, resulting in a total of 54 covariates for model training. Twelve models were developed for the six soil properties, both with and without CrFr covariates. Validation results showed that including CrFr covariates improved the accuracy of models for BD, OC, AT, and TN. However, the impact on models for pH and CEC was minimal, indicating that intrinsic soil factors likely influence these properties more than CrFr. Partial dependence plots indicated that the models captured expected patterns, such as the negative association of forest cover with BD and its positive relationship with OC and TN. In contrast, crops such as legumes and corn exhibit the opposite effects. Forests exhibited a negative relationship with AT, whereas croplands showed a positive association, indicating a likely difference between the Ap horizon and Ah. Uncertainty analysis revealed lower uncertainty in agricultural cropland areas and those with lower elevations. This study highlights the potential of DSM in assessing the impact of crop type on soils and suggesting what crops may be more beneficial for soil.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".