AI-vergreen: a multi-label Sentinel-2 training dataset of summer green (Larix) and evergreen needle leaf forest types in boreal forest biomes for remote sensing applications
Bibliographic record
Abstract
Boreal forests, which represent roughly one-third of the world’s total forested area, provide critical ecosystem services including carbon stocks, climate feedback, permafrost stability, biodiversity, and economic benefits. Located in the northern latitude, they are mainly dominated by evergreen needle-leaf tree taxa (Pinus, Picea, Abies) in North America, Northern Europe, and Western Siberia, and by deciduous needle-leaf tree taxa (Larix) in Eastern Siberia. Remote sensing applications in high latitudes are possible but remain challenging for optical satellite sensors due to frequent cloud coverage, forest fires, and low illumination. Additionally, there is little data available prepared as multi-label datasets for remote sensing applications focusing on the structure of boreal forests, specifically on Larix deciduous trees. Furthermore, labeled datasets of summer green and evergreen forest types for specific satellite sensors would enable remote sensing and deep learning applications such as classification, and ultimately improve our understanding of evergreen and summer green tree dynamics. An example of such a dataset is the TreeSatAI multi-sensor Artificial Intelligence Benchmark Archive (doi.org/10.5281/zenodo.6780578), which provides labels on species and forest composition in Europe. Another one is the SiDroForest data collection, consisting of a synthetic Unmanned Aerial Vehicle (UAV) Siberian Larch Dataset (doi.org/10.1594/PANGAEA.932795) and Sentinel-2 image patches (doi.org/10.1594/PANGAEA.933268) of 54 forest plots in Eastern Siberia. Here we are building up an extensive multi-labeled training dataset based on optical Sentinel-2 image patches (60 x 60 m image patch of the 10 m and 20 m S2-bands), including meta-data information on summer green and evergreen tree species and forest structure from vegetation plots. Over 250 vegetation plots were collected since 2011 from nine field expeditions of the Alfred Wegener Institute in Eastern Siberia (doi.org/10.5194/essd-14-5695-2022) and Western Canada, where vegetation was sampled and described, and UAV images were taken (UAV solely in 2021 and 2022). In addition to in-situ plots, we gathered all cloud-free Sentinel-2 data from late spring to early fall (May to October) that geographically coincides with the vegetation plots. Therefore, the dataset contains different phenophases of evergreen and summer green forests and provides detailed label information on forest structure – such as tree species and density. The multi-labeling will include broader and more detailed forest-type classes. Some examples of higher-level labels are “Sparse larch forest” or “Dense evergreen forest’’. The poster will demonstrate how we defined forest labels from in-situ data, UAV, Sentinel-2, and their corresponding spectral signatures.We anticipate our dataset to be a starting point for a significantly more extensive one with the addition of radar satellite sensors such as Sentinel-1 and TanDEM-X, and other ground vegetation plots (new expedition expected in Alaska and Canada in summer 2023), data search in literature and repositories– e.g. NASA Arctic Boreal Vulnerability Experiment. Our dataset will be publicly available and can be used as a training dataset for deep learning algorithms to identify and characterize evergreen and summer green needle-leaf trees in boreal forest regions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.003 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".