Species composition, forest properties and land cover types across Canada’s forests at 250m resolution for 2001 and 2011
Bibliographic record
Abstract
This data publication contains two collections of raster maps of forest attributes across Canada, the first collection for year 2001, and the second for year 2011. The 2001 collection is actually an improved version of an earlier set of maps produced also for year 2001 (Beaudoin et al 2014, DOI: https://doi.org/10.1139/cjfr-2013-0401) that is itself available through the web site “http://nfi-nfis.org”. Each collection contains 93 maps of forest attributes: four land cover classes, 11 continuous stand-level structure variables such as age, volume, biomass and height, and 78 continuous values of percent composition for tree species or genus. The mapping was done at a spatial resolution of 250m along the MODIS grid. Briefly the method uses forest polygon information from the first version of photoplots database from Canada’s National Forest Inventory as reference data, and the non-parametric k-nearest neighbors procedure (kNN) to create the raster maps of forest attributes. The approach uses a set of 20 predictive variables that include MODIS spectral reflectance data, as well as topographic and climate data. Estimates are carried out on target pixels across all Canada treed landmass that are stratified as either forest or non-forest with 25% forest cover used as a threshold. Forest cover information was extracted from the global forest cover product of Hansen et al (2013) (DOI: https://doi.org/10.1126/science.1244693). The mapping methodology and resultant datasets were intended to address the discontinuities across provincial borders created by their large differences in forest inventory standards. Analysis of residuals has failed to reveal residual discontinuities across provincial boundaries in the current raster dataset, meaning that our goal of providing discontinuity-free maps has been reached. The dataset was developed specifically to address strategic issues related to phenomena that span multiple provinces such as fire risk, insect spread and drought. In addition, the use of the kNN approach results in the maintenance of a realistic covariance structure among the different variable maps, an important property when the data are extracted to be used in models of ecosystem processes. For example, within each pixel, the composition values of all tree species add to 100%. A scientific publication is under review, the reference will be provided.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.005 | 0.012 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.010 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".