Regional-scale digital soil mapping in british columbia using legacy soil survey data and machine-learning techniques
Bibliographic record
Abstract
Digital soil mapping (DSM) is the intersection of geographical information systems (GIS), and (spatial) statistics and is a sub-discipline of soil science that has been increasingly relevant in helping to address emerging issues such as food production, climate change, land resource management, and the management of earth systems. Even with the need for digital soil information in the raster format, such information is limited for British Columbia (BC) where much of it is digitized from legacy soil survey maps with inherent spatial problems related to polygon boundaries; attribute specificity due to multi-component map units; and map scale where small-scale surveys have limited use in addressing local and regional needs. In spite of these issues, legacy soil survey data are still useful as sources of training data where machine-learning techniques may be used to extract soil-environmental relationships from a survey and a suite of digital environmental covariates. This dissertation describes a framework for developing training data from conventional soil survey maps and compares various machine-learning techniques for predicting the spatial patterns of qualitative soil data such as soil parent material and soil classes. Results of this research included maps of soil parent material, Great Groups, and Orders for the Lower Fraser Valley and a soil Great Group map for the Okanagan-Kamloops region at a 100 m spatial resolution. Key findings included (1) the recognition of Random Forest being the most effective machine-learner based on two model comparison studies; (2) the conclusion that model choice greatly impacted the accuracy of predictions; (3) the method for developing training data greatly impacted the accuracy through a comparison of four methods; and (4) that training data derived from soil survey maps were more effective in representing the feature space of various classes in comparison to using training data derived from soil pits. This study advances the understanding of model selection and training data development in DSM and may facilitate the future development of methodologies for provincial maps of BC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".