MétaCan
Menu
← Back to cohort
Record W4322006239 · doi:10.5194/egusphere-egu23-7624

AI-vergreen: a multi-label Sentinel-2 training dataset of summer green (Larix) and evergreen needle leaf forest types in boreal forest biomes for remote sensing applications

2023· preprint· en· W4322006239 on OpenAlexaboutno aff
Léa Enguehard, Birgit Heim, Stefan Kruse, Begüm Demir, Robert Jackisch, Josias Gloy, Sarah Haupt, Laura Schild, Femke van Geffen, Veronika Döpper, Ronny Hänsch, Nicola Falco, Ulrike Herzschuh

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldEnvironmental Science
TopicRemote Sensing and LiDAR Applications
Canadian institutionsnot available
Fundersnot available
KeywordsEvergreenBiomeDeciduousRemote sensingTaigaLarchBorealEnvironmental scienceGeographyPhysical geographyForestryEcosystemEcologyBiology

Abstract

fetched live from OpenAlex

Boreal forests, which represent roughly one-third of the world’s total forested area, provide critical ecosystem services including carbon stocks, climate feedback, permafrost stability, biodiversity, and economic benefits. Located in the northern latitude, they are mainly dominated by evergreen needle-leaf tree taxa (Pinus, Picea, Abies) in North America, Northern Europe, and Western Siberia, and by deciduous needle-leaf tree taxa (Larix) in Eastern Siberia. Remote sensing applications in high latitudes are possible but remain challenging for optical satellite sensors due to frequent cloud coverage, forest fires, and low illumination. Additionally, there is little data available prepared as multi-label datasets for remote sensing applications focusing on the structure of boreal forests, specifically on Larix deciduous trees. Furthermore, labeled datasets of summer green and evergreen forest types for specific satellite sensors would enable remote sensing and deep learning applications such as classification, and ultimately improve our understanding of evergreen and summer green tree dynamics. An example of such a dataset is the TreeSatAI multi-sensor Artificial Intelligence Benchmark Archive (doi.org/10.5281/zenodo.6780578), which provides labels on species and forest composition in Europe. Another one is the SiDroForest data collection, consisting of a synthetic Unmanned Aerial Vehicle (UAV) Siberian Larch Dataset (doi.org/10.1594/PANGAEA.932795) and Sentinel-2 image patches (doi.org/10.1594/PANGAEA.933268) of 54 forest plots in Eastern Siberia. Here we are building up an extensive multi-labeled training dataset based on optical Sentinel-2 image patches (60 x 60 m image patch of the 10 m and 20 m S2-bands), including meta-data information on summer green and evergreen tree species and forest structure from vegetation plots. Over 250 vegetation plots were collected since 2011 from nine field expeditions of the Alfred Wegener Institute in Eastern Siberia (doi.org/10.5194/essd-14-5695-2022) and Western Canada, where vegetation was sampled and described, and UAV images were taken (UAV solely in 2021 and 2022). In addition to in-situ plots, we gathered all cloud-free Sentinel-2 data from late spring to early fall (May to October) that geographically coincides with the vegetation plots. Therefore, the dataset contains different phenophases of evergreen and summer green forests and provides detailed label information on forest structure – such as tree species and density. The multi-labeling will include broader and more detailed forest-type classes. Some examples of higher-level labels are “Sparse larch forest” or “Dense evergreen forest’’. The poster will demonstrate how we defined forest labels from in-situ data, UAV, Sentinel-2, and their corresponding spectral signatures.We anticipate our dataset to be a starting point for a significantly more extensive one with the addition of radar satellite sensors such as Sentinel-1 and TanDEM-X, and other ground vegetation plots (new expedition expected in Alaska and Canada in summer 2023), data search in literature and repositories– e.g. NASA Arctic Boreal Vulnerability Experiment. Our dataset will be publicly available and can be used as a training dataset for deep learning algorithms to identify and characterize evergreen and summer green needle-leaf trees in boreal forest regions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.047
Threshold uncertainty score0.093

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0030.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0030.001
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0040.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.074
GPT teacher head0.320
Teacher spread0.246 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same topicRemote Sensing and LiDAR Applications→French-language works237,207→