Comparison of Models for Spatial Distribution and Prediction of Cadmium in Subtropical Forest Soils, Guangdong, China
Bibliographic record
Abstract
Cadmium (Cd) is a toxic metal and found in various soils, including forest soils. The great spatial heterogeneity in soil Cd makes it difficult to determine its distribution. Both traditional soil surveys and spatial modeling have been used to study the natural distribution of Cd. However, traditional methods are highly labor-intensive and expensive, while modeling is often encumbered by the need to select the proper predictors. In this study, based on intensive soil sampling (385 soil pits plus 64 verification soil pits) in subtropical forests in Yunfu, Guangdong, China, we examined the impacting factors and the possibility of combining existing soil information with digital elevation model (DEM)-derived variables to predict the Cd concentration at different soil depths along the landscape. A well-developed artificial neural network model (ANN), multi-variate analysis, and principal component analysis were used and compared using the same dataset. The results show that soil Cd concentration varied with soil depth and was affected by the top 0–20 cm soil properties, such as soil sand or clay content, and some DEM-related variables (e.g., slope and vertical slope position, varying with depth). The vertical variability in Cd content was found to be correlated with metal contents (e.g., Cu, Zn, Pb, Ni) and Cd contents in the layer immediately above. The selection of candidate predictors differed among different prediction models. The ANN models showed acceptable accuracy (around 30% of predictions have a relative error of less than 10%) and could be used to assess the large-scale Cd impact on environmental quality in the context of intensifying industrialization and climate change, particularly for ecosystem management in this region or other regions with similar conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".