Machine Learning-Based Mapping of Lake Ice Cover from SWOT KaRIn Backscatter: Preliminary Results
Notice bibliographique
Résumé
Lakes are key components of the global freshwater system, playing a crucial role in climate regulation, hydrological cycling, and maintaining ecological balance. As highly sensitive indicators of climate change, lakes are recognized by the Global Climate Observing System (GCOS) as an essential climate variable (ECV), with lake ice cover (LIC) and lake ice thickness (LIT) identified as two of its thematic products. In the Northern Hemisphere, many lakes develop seasonal ice cover, which significantly influences local energy balance, ecosystem function, and socio-economic activities such as transportation, fishing, recreation, and tourism. Understanding the spatial distribution and temporal dynamics of the lake surface conditions is essential for numerous applications. For instance, accurate mapping of ice cover dynamics in lakes is crucial for predicting lake ice phenology, estimating ice thickness, and assessing the impacts of climate change on lake ecosystems. Due to the steady decline in long-term in situ observations of lake ice and overlying snow properties in recent decades, there is an increasing reliance on satellite remote sensing. These spaceborne approaches provide an effective alternative for investigating lakes at regional and global scales, providing a comprehensive and cost-effective means of monitoring these dynamic water bodies. The Surface Water and Ocean Topography (SWOT) mission, launched in December 2022, introduces the Ka-band Radar Interferometer (KaRIn), which offers high-resolution measurements of both surface water elevation and backscatter. While primarily designed for hydrological and oceanographic studies, SWOT's KaRIn backscatter data have shown sensitivity to surface conditions, suggesting potential applications in cryospheric monitoring. Our analyses demonstrate that KaRIn backscatter is sensitive to lake surface conditions and exhibits a clear distinction between ice-covered and open-water areas, providing a basis for classification efforts. Building on these initial findings, we develop a Random Forest classification model to distinguish between lake ice and open water using SWOT KaRIn backscatter data. Random Forest is a robust machine learning algorithm known for its effectiveness in handling complex, nonlinear relationships and has been widely used in remote sensing classification studies. The study focuses on five lakes in the Northern Hemisphere: Lake Teshekpuk in Alaska, as well as Kluane Lake, Great Slave Lake, Lake Athabasca, and Great Bear Lake, all located in Canada. Lake Teshekpuk and Kluane Lake are consistently observed during both the Cal/Val and Science phases of the SWOT mission, while Great Slave Lake, Lake Athabasca, and Great Bear Lake were partially covered during the Cal/Val phase, with full spatial coverage during the Science phase. The time period considered for performing the classification spans from March 30 to July 10, 2023 (Cal/Val phase), and from September 1, 2023, to July 15, 2025 (Science phase). To support classification and provide accurate labelling, we use supplementary satellite datasets including Sentinel-1 SAR, Sentinel-2 MSI, MODIS Aqua/Terra, and Landsat 8/9. These datasets provide additional information on surface conditions to ensure accurate reference labelling. By leveraging SWOT's high-resolution data and the capabilities of machine learning, this study aims to enhance the monitoring of lake ice phenology. It offers valuable insights into the value of Ka-band for mapping lake ice.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».