Bibliographic record
Abstract
The China-Pakistan Economic Corridor (CPEC) is one of flagship projects of the One Belt One Road Initiative, which faces threats from mountain disasters in the high altitude region, such as glacial lake outburst floods (GLOFs). An up-to-date high-quality glacial lake dataset with critical parameters (e.g. lake types), which is fundamental to flood risk assessments and predicting glacier-lake evolutions, is still largely absent for the entire CPEC. This study describes a glacial lake dataset in 2020 for CPEC at 10–30 m resolution, which was produced from both Landsat and Sentinel optical images as well as glacial lake inventories in 1990 and 2000 from Landsat observation, using an advanced object-oriented mapping method associated with rigorous visual inspection workflows. The results show that Landsat derived 2234 glacial lakes in 2020, covering a total area of 86.31 ± 14.98 km2 with a minimum mapping unit of 5 pixels (4500 m2), whereas Sentinel derived 7560 glacial lakes in 2020 with a total area of 103.70 ± 8.45 km2 with a minimum mapping unit of 5 pixels (500 m2). The discrepancy implies that there is a significant quantity of small glacier lakes not recognized in existing glacial lake inventories and a more thorough inclusion of them require future efforts using higher resolution data. The total number and area of glacial lakes from consistent 30 m resolution Landsat images remain relatively stable despite a slight increase from 1990 to 2020. A range of critical attributes have been generated in the dataset, including lake types of two classification systems and mapping uncertainty estimated by an improved equation. This comprehensive glacial lake dataset has potentials to be widely applied in studies on glacial lake-related hazards and glacier-lake interactions, and is freely available at https://doi.org/10.12380/Glaci.msdc.000001 (Lesi et al., 2022).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.032 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.025 | 0.025 |
| Insufficient payload (model declined to judge) | 0.128 | 0.091 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".