SiDroForest  Siberian Drone-mapped Forest inventory
Bibliographic record
Abstract
To gain a better understanding of global carbon storage and albedo feedback mechanisms it is important to have insights into high latitude vegetation change. Boreal forest compositions are changing in response to changes in climate, which in turn can lead to feedbacks in regional and global climate through altered carbon cycles and albedo dynamics. Circumpolar boreal forests represent close to 30% of all forested area on the planet, between 900 and 1,200 million ha. These forests are located primarily in Alaska, Canada, and Russia. Due to the remote location of these forests and the short seasons without snow, data collected on the boreal vegetation is limited. The proposed dataset is an attempt to remedy data scarcity whilst providing adjusted data for machine learning practices.We present a dataset containing diverse formats of forest structure information that covers two important vegetation transition zones in Siberia: the Evergreen - Summergreen transition zone in Central Yakutia and the northern treeline in Chukotka (NE Siberia). This dataset contains data from the locations covered by fieldwork was performed by the Alfred Wegener Institute for Polar and Marine research, (AWI) and the North-Eastern Federal University of Yakutsk (NEFU). The fieldwork upscaled through the addition of Red Green Blue(RGB) UAV (Unmanned Aerial Vehicle) camera data and Sentinel-2 satellite data cropped to a 5 km radius around the fieldwork sites. The dataset is created with the aim of providing ground truth validation and training data to be used in various vegetation related machine learning tasks . The dataset contains: 1.Labelled individual trees per 30x30 m plot assigned in field work with additional data on species, height, crown width, and biomass. 2.Structure from Motion (SfM)point clouds that provide 3D information about the forest structure, included generated Canopy Height Model (CHM), Digital Elevation Model (DEM) and a Digital Surface Model (DSM) per 50x50 m. 3.Multispectral Sentinel-2 satellite data (10 m ) cropped to a 5km radius with generated a NDVI(normalized difference vegetation index), available in three seasons: Early Summer, Peak Summer and Late Summer. 4.Extracted tree crowns with species information and a synthetically generated large (10.000 samples) dataset for training machine leaning algorithms. The dataset will be made publicly available on the data repository PANGAEA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.016 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".