Erosion-SAM: Semantic segmentation of soil erosion by water
Bibliographic record
Abstract
• Fine-tuned SAM improves soil erosion segmentation accuracy over the original model. • Segmentation accuracy is highest in grassland compared to other land covers. • Prompt-based resizing enhances erosion detection across all land cover types. • Fine-tuned SAM provides high-quality training data for erosion prediction models. Soil erosion (SE) by water threatens global agriculture by depleting fertile topsoil and causing economic costs. Conventional SE models struggle to capture the complex, non-linear interactions between SE drivers. Recently, machine learning has gained attention for SE modeling. However, machine learning requires large data sets for effective training and validation. In this study, we present Erosion-SAM, which fine-tunes the Segment Anything Model (SAM) for automatic segmentation of water erosion features in high-resolution remote sensing imagery. The data set comprised 405 manually segmented agricultural fields from erosion-prone areas obtained from the rain gauge-adjusted radar rainfall data (RADOLAN) for bare cropland, vegetated cropland, and grassland. Three approaches were evaluated: two pre-processing techniques— resizing and cropping — and an improved version of the resizing approach with user-defined prompts during the testing phase. All fine-tuned models outperformed the original SAM, with the prompt-based resizing method showing the highest accuracy, especially for grassland (recall: 0.90, precision: 0.82, dice coefficient: 0.86, IoU: 0.75). SAM performed better than the cropping approach only on bare cropland. This discrepancy is attributed to the tendency of SAM to overestimate SE by classifying a large proportion of fields as eroded, which increases recall by covering most of the eroded pixels. All three fine-tuned approaches showed strong correlations with the actual SE severity ratios, with the prompt-enhanced resizing approach achieving the highest R 2 of 0.93. In summary, Erosion-SAM shows promising potential for automatically detecting SE features from remote sensing images. The generated data sets can be applied to machine learning-based SE modeling, providing accurate and consistent training data across different land cover types, and offering a reliable alternative to traditional SE models. In addition, erosion-SAM can make a valuable contribution to the precise monitoring of SE with high temporal resolution over large areas, and its results could benefit reinsurance and insurance-related risk solutions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".