Soil data for mapping paludification in black spruce forests of eastern Canada
Bibliographic record
Abstract
Soil data and soil mapping are indispensable tools in sustainable forest management. In northern boreal ecosystems, paludification is defined as the accumulation of partially decomposed organic matter over saturated mineral soils, a process that reduces tree regeneration and forest growth. Given this negative effect on forest productivity, spatial prediction of paludification in black spruce stands is important in forest management. This paper provides a description of the soil database to predict organic layer thickness (OLT) as a proxy of paludification in northeastern Canada. The database contains 13,944 OLT measurements (in cm) and their respective GPS coordinates. We collected OLT measurements from georeferenced ground plots and transects from several previous projects. Despite the variety of sources, the sampling design for each dataset was similar, consisting of manual measurements of OLT with a hand probe. OLT measurements were variable across the study area, with a mean ± standard deviation of 21 ± 24 cm (ranging from a minimum of 0 cm to a maximum of 150 cm), and the distribution tended toward positive skewing, with a large number of low OLT values and fewer high OLT values. The dataset has been used to perform OLT mapping at 30-m resolution and predict the risk of paludification in northeastern Canada (Mansuy et al., 2018) [1]. The spatially explicit and continuous database is also available to support national and international efforts in digital soil mapping.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.008 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".