Quantifying the influencing factors and multi-factor interactions affecting cadmium accumulation in limestone-derived agricultural soil using random forest (RF) approach
Bibliographic record
Abstract
Cadmium (Cd) is a highly toxic heavy metal that occurs widely in the environment and poses extensive threats to human health, animals, and plants. This study aims to identify and apportion multi-source and multi-phase Cd pollution from natural and anthropogenic inputs using ensemble models that include random forest (RF) in agricultural soils on Karst areas. The contributions of natural and anthropogenic factors to Cd accumulation were quantitatively assessed using the RF machine learning method. The results revealed that the main influencing factors were pH, organic carbon (Corg), and elevation. Moreover, the interaction effects of pH and Corg on distance and elevation were also quantified and visualised. It is observed that pH and Corg had stronger effects on soil Cd concentration than that of distance when pH > 7.02 and Corg > 1.53. In other words, higher Cd content in the soil along roadways may be caused by the interaction of distance, pH and Corg, with pH and Corg playing the dominant role in our case. Moreover, the maximum contribution of a single factor, elevation, to Cd concentration was about 0.13 mg/kg, and its interactions reached 1.082 mg/kg and 0.83 mg/kg, respectively, when combined with pH and Corg at 194.0 m. However, with increasing elevation, pH and Corg gradually took over the leading roles. This result not only gives us a quantitative understanding of the relationship between the factors that affect soil cadmium accumulation, but also provides an accurate method for source apportionment of heavy metals in soil.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".