AI-vergreen: a multi-label Sentinel-2 training dataset of summer green (Larix) and evergreen needle leaf forest types in boreal forest biomes for remote sensing applications
Notice bibliographique
Résumé
Boreal forests, which represent roughly one-third of the world’s total forested area, provide critical ecosystem services including carbon stocks, climate feedback, permafrost stability, biodiversity, and economic benefits. Located in the northern latitude, they are mainly dominated by evergreen needle-leaf tree taxa (Pinus, Picea, Abies) in North America, Northern Europe, and Western Siberia, and by deciduous needle-leaf tree taxa (Larix) in Eastern Siberia. Remote sensing applications in high latitudes are possible but remain challenging for optical satellite sensors due to frequent cloud coverage, forest fires, and low illumination. Additionally, there is little data available prepared as multi-label datasets for remote sensing applications focusing on the structure of boreal forests, specifically on Larix deciduous trees. Furthermore, labeled datasets of summer green and evergreen forest types for specific satellite sensors would enable remote sensing and deep learning applications such as classification, and ultimately improve our understanding of evergreen and summer green tree dynamics. An example of such a dataset is the TreeSatAI multi-sensor Artificial Intelligence Benchmark Archive (doi.org/10.5281/zenodo.6780578), which provides labels on species and forest composition in Europe. Another one is the SiDroForest data collection, consisting of a synthetic Unmanned Aerial Vehicle (UAV) Siberian Larch Dataset (doi.org/10.1594/PANGAEA.932795) and Sentinel-2 image patches (doi.org/10.1594/PANGAEA.933268) of 54 forest plots in Eastern Siberia. Here we are building up an extensive multi-labeled training dataset based on optical Sentinel-2 image patches (60 x 60 m image patch of the 10 m and 20 m S2-bands), including meta-data information on summer green and evergreen tree species and forest structure from vegetation plots. Over 250 vegetation plots were collected since 2011 from nine field expeditions of the Alfred Wegener Institute in Eastern Siberia (doi.org/10.5194/essd-14-5695-2022) and Western Canada, where vegetation was sampled and described, and UAV images were taken (UAV solely in 2021 and 2022). In addition to in-situ plots, we gathered all cloud-free Sentinel-2 data from late spring to early fall (May to October) that geographically coincides with the vegetation plots. Therefore, the dataset contains different phenophases of evergreen and summer green forests and provides detailed label information on forest structure – such as tree species and density. The multi-labeling will include broader and more detailed forest-type classes. Some examples of higher-level labels are “Sparse larch forest” or “Dense evergreen forest’’. The poster will demonstrate how we defined forest labels from in-situ data, UAV, Sentinel-2, and their corresponding spectral signatures.We anticipate our dataset to be a starting point for a significantly more extensive one with the addition of radar satellite sensors such as Sentinel-1 and TanDEM-X, and other ground vegetation plots (new expedition expected in Alaska and Canada in summer 2023), data search in literature and repositories– e.g. NASA Arctic Boreal Vulnerability Experiment. Our dataset will be publicly available and can be used as a training dataset for deep learning algorithms to identify and characterize evergreen and summer green needle-leaf trees in boreal forest regions.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,003 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,003 | 0,001 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,006 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».