Assessing machine learning models to generate permafrost distribution map in Solukhumbu, Nepal
Bibliographic record
Abstract
Permafrost is one of the key components of the cryosphere. Previous studies show that the extent of permafrost has shifted to higher elevations in Nepal. These researches, however, has been hampered by inconsistency in their study period. Proxies like rock glaciers and climatic variables, such as multi-decadal annual air temperature, are used to link towards the likely occurrence of permafrost. Here, the rock glacier inventory of Solukhumbu was prepared, and classified based on their activity (Intact/Relict) from Google Earth. Talus-based rock glaciers were observed more than glacier-derived ones. These rock glaciers were highly correlated with Mean Annual Air Temperature, followed by potential incoming solar radiation and slope. Three machine learning models (Logistic Regression, Random Forest and Support Vector Machines) were trained to generate permafrost probability distribution maps based on their prediction of the probability of rock glaciers being intact as opposed to relict. Logistic Regression and Support Vector Machines were able to produce a similar spatial distribution of permafrost. However, the Random Forest has low precision of spatial variation. The permafrost distribution map suggests the likely occurrence of permafrost to be above 5000 m, indicating a potential for rock and landslides should it thaw in the future. While higher-resolution input data can improve the results, this approach remains promising for application in permafrost regions where information about the ice content of rock glaciers is very limited.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".