Provincial-scale digital soil mapping using a random forest approach for British Columbia
Bibliographic record
Abstract
Although British Columbia (BC), Canada, has a rich history of producing conventional soil maps (CSMs) between 1925 and 2000, the province still lacks a detailed soil map with a comprehensive coverage due to the cost and time required to develop such a product. This study builds on previous digital soil mapping (DSM) research in BC and develops provincial-scale maps. Soil taxonomic classes (e.g., great groups and order) and parent material classes were mapped at a 100 m spatial resolution for BC (944 735 km 2 ). Training points were generated from detailed and semi-detailed soil survey maps. The training points were intersected with 26 topographic indices for mapping parent materials with an additional 9 climatic and vegetation indices for mapping soil classes. The soil–environmental relationships were inferred using the random forest (RF) classifier. The fitted models were used to predict 23 soil great groups, 9 soil orders, and 10 parent material classes. Accuracy assessments were performed using n = 14 570 validation points for parent material classes and n = 14 316 validation points for soil classes, acquired from the BC Soil Information System. The accuracy rates for soil great groups, orders, and parent material classes were 55%, 62%, and 69%, respectively, and kappa coefficients were 0.37, 0.41, and 0.59, respectively. This study demonstrated that when RF was trained using CSMs, the accuracy for the resulting DSM was higher than the original CSM. To assess prediction uncertainties, ignorance uncertainty maps were developed using class-probability layers generated by the RF models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".