Prediction of Pb and Zn in urban soil using VIS-NIR-SWIR spectroscopy
Bibliographic record
Abstract
Heavy metals serve as a subset of chemical elements with higher density than iron. Besides, these environmental pollutants are constant and nonbiodegradable elements that can cause toxicity and genetic mutations to the live cells. Depending on the study area, an increase in soil heavy metals from a specific level often created by human activities can lead to many adverse effects on individuals, soil, and plants. In case of their existence in the food chain or transfer to groundwater resources, human health is seriously threatened. Over numerous years, being affected by a colossal number of pollutant resources such as world war and household waste, industry, transportation systems, and urbanization has changed Berlin to a city at risk of soil pollution by heavy metals. That is why carrying out a study on heavy metals in this city is of great significance. Chemical analysis is the first and most traditional ways to measure soil heavy metals. Despite high precision, this method is complicated, time-consuming, costly, and ineffective on a large scale. However, the spectral data facilitates the rapid and cost-effective assessment of these elements. Therefore, in this study, the ability of spectral data to predict heavy metals in Berlin’s soil is examined.When it comes to the data required, there are two categories: 1) heavy metals (Pb and Zn) related to more than 600 soil samples collected from 2016 to 2018 and measured in the laboratory, and 2) the spectral data measured for each sample in the range between 350 to 2500nm in a spectrometry lab. All data is divided into training (80%) and testing (20%) to reach this aim. Next, the first group is used to train the machine learning algorithms, including partial least square regression (PLSR), support vector regression (SVR), and random forest (RF). Moreover, the second group is used to test the models. Finally, the accuracy of models is evaluated by correlation of determination (R2), and Root mean square error (MSE). As a part of the results, R2 and MSE were achieved 0.25, and 4394.45 for Pb, and 0.18 and 6558.49 for Zn.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".