Multi-Temporal Landsat-8 Images for Retrieval and Broad Scale Mapping of Soil Copper Concentration Using Empirical Models
Bibliographic record
Abstract
Mapping soil heavy metal concentration using machine learning models based on readily available satellite remote sensing images is highly desirable. Accurate mapping relies on appropriate data, feature extraction, and model selection. To this end, a data processing pipeline for soil copper (Cu) concentration estimation has been designed. First, instead of using single Landsat scenes, the utilization of multiple Landsat scenes of the same location over time is considered. Second, to generate a preferred feature set as input to a regression model, a number of feature extraction methods are motivated and compared. Third, to find a preferred regression model, a variety of approaches are implemented and compared for accuracy. In this research, 11 Landsat-8 images from 2013 to 2017 of Gulin County, Sichuan China, and 138 soil samples with lab-measured Cu concentrations collected from the area in 2015 are used. A variety a metrics under cross-validation are used for comparison. The results indicate that multi-temporal images increase accuracy compared to single Landsat images. The preferred feature extraction varies based on the regression model used; however, the best results are obtained using support vector regression and the original data. The final soil Cu map generated using the recommended data processing pipeline shows a consistent spatial pattern with a ground-truth land cover classification map. These results indicate that machine learning has the ability to perform large-scale soil heavy metal mapping from widely available satellite remote sensing images.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".