Deep Learning for mapping retrogressive thaw slumps across the Arctic
Bibliographic record
Abstract
Retrogressive thaw slumps (RTS) are typical landscape processes of thawing and degrading permafrost. To this point, their distribution and dynamics are almost completely undocumented across many regions in the permafrost domain, partially due to the lack of data and monitoring techniques in the past. We are tackling this shortcoming by creating a deep learning based semantic segmentation framework to detect RTS, using multi-spectral PlanetScope, derived topographic (ArcticDEM) and multi-temporal Landsat Trend data. We created a highly automated processing pipeline, which is designed to create reproducible results and to be flexible for multiple input features. The processing workflow is based on the pytorch deep-learning framework and includes a variety of different segmentation architectures (UNet, UNet++, DeepLabV3), backbones and includes common data transformation techniques such as augmentation or normalization. \nWe tested (training, validation) our DL based model in six different regions of 100 to 300 km² size across Canada (Banks Island, Tuktoyaktuk, Horton, Herschel Is.), and Siberia (Kolguev, Lena). We performed a regional cross-validation (5 regions training, 1 region validation) to test the spatial robustness and transferability of the algorithm. Furthermore, we tested different architectures backbones and loss-function to identify the best performing and most robust parameter sets. For training the models we created a training database of manually digitized and validated RTS polygons. \nThe resulting model performance varied strongly between different regions with maximum Intersection over Union (IoU) scores between 0.15 and 0.58. The strong regional variation emphasizes the need for sufficiently large training data, which is representative for the massive variety of RTS. However, the creation of good training data proved to be challenging due to the fuzzy definition and delineation of RTS, particularly on the lower part. \nWe have recently expanded our analysis to several RTS-rich regions across the Arctic (Fig.X) for the year 2021 and annual analysis (2018-2021) for RTS hot-spots, e.g. Banks Island, Peel Plateau and others. First model inference runs are promising for detecting RTS, but are still strongly overestimating the number and area of RTS, due to an excessive number of false positives. Model performance however, varies strongly between regions. Due to the strong variability of landscapes with RTS, we expect an improvement in model performance with an increase in the number and spatial distribution of training datasets. The community driven formation of the IPA Action Group RTSIn, which aims to create standardized RTS digitization protocols and training datasets for deep/machine-learning purposes will be a great boost for our purpose. \nWith our standardized processing pipeline (preprocessing, training, inference), which allows to add more features based on user interest and data availability,, we tested our workflow for surface water and pingos with a mixture of publically available (Jones et al) and digitized data (Grosse pingos, Nitze water). These tests produced very good results and showed that the designed workflow is transferrable beyond the segmentation of RTS only. \nIn the near future, we are aiming to integrate the community based training data and further gradually improve our training database. Within the framework of the ML4Earth project, we will create a temporal and pan-arctic monitoring system for RTS based on our highly automated processing chain. This will enable us to better understand pan-arctic RTS dynamics, their influencing factors, and consequences. Combining these spatial-temporal datasets with volumetric change information and carbon stock information will enable us to better quantify the consequences of thaw slumping across the permafrost domain.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".