Spatial downscaling of global soil texture classes into 30 m images at the province scale
Bibliographic record
Abstract
Soil categorical data is an important aspect in soil science because it effectively facilitates communication between policymakers and stakeholders. Furthermore, soil categorical data exceeds single-property data in terms of depth of information and is an essential component of various scientific disciplines such as hydrological, ecological and pollution frameworks. However, datasets containing such information are usually too coarse for local needs and regional policies . In this study, an algorithm was introduced, known as rafikisol , to spatially downscale (to a finer detail/resolution) soil texture classes from 1 km SoilGrids images into 30 m images in three environmentally diverse provinces (Gauteng, KwaZulu-Natal, and the Western Cape) in South Africa . Rafikisol surpassed the performance of another high-resolution soil dataset (Innovative Solutions for Decision Agriculture) by 9% and 27% in Gauteng and the Western Cape, respectively (accuracy ∼75% and ∼72%). Conversely, iSDAsoil outperformed rafikisol by 34% in KwaZulu-Natal (11% accuracy). The spatial soil texture class distribution predicted by rafikisol was considerably different and heavier (clayey) in Gauteng and KwaZulu-Natal but had a similar spatial distribution with lighter (sandy) soil texture in the Western Cape compared to iSDAsoil. With improvements such as the introduction of new sampling techniques, algorithm optimization and the use of expert knowledge, this method has the potential to increase the accuracy of additional modeling frameworks that require high-resolution soil information, especially in data-scarce or resource-constrained regions. This has implications for proper land-use management, affecting aspects ranging from food security and urban expansion to biodiversity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".