Lidar plots and point cloud metrics derived from airborne laser scanning transects acquired over forests in northern Canada.
Bibliographic record
Abstract
The datasets consist of two relational databases containing lidar plots and point cloud metrics derived from airborne laser scanning (ALS) transects. The ALS data were acquired in 2023 over forest-dominated ecozones in northern Canada. The transects had a minimum swath width of 500 m, total ~20,000 km in length, and include ~15 million lidar plots. The datasets are delivered as SQLite GeoPacakages, with a full version including 369 point cloud metrics and an abridged version including 40 metrics. Data for additional years and detailed documentation will be made available through Canada’s National Forest Information System: https://opendata.nfis.org/mapserver/nfis-change_eng.html A thorough description of the dataset and comparison with satellite information products can be found in the accompanying manuscript: Bater, C.W., White, J.C., Chen, H., Tompalski, P., Hermosilla, T., Boucher, J., Wulder, M.A. Submitted. Airborne laser scanning transects over Canada’s northern forests: lidar plots for science and application. Earth System Science Data. Abstract (from the manuscript): Mapping vegetation is required for monitoring the condition of forest resources. Satellite data provide information on land cover and change; however, forest structural attributes are difficult to model without additional measurements from ground plots or airborne laser scanning (ALS, also known as airborne light detection and ranging or lidar) instruments. Over large and inaccessible areas, such as Canada’s northern and predominantly unmanaged forests, ground plots are expensive, difficult to install, and unlikely to form a statistically valid probability sample. An alternative means to obtain information regarding forest structure in these situations is samples of ALS (hereafter lidar plots). Transect-based samples of ALS data can be used to provide structural information for the calibration and validation of spatially explicit predictive modelling for wide-area mapping of forest attributes. Here we describe and share data from the recent acquisition and processing of ALS transects across Canada’s northern forests. To date, approximately 43,000 km of ALS transects have been acquired in 2023 and 2024, with additional coverage ongoing for 2025. Acquisition flight lines were designed to sample a range of northern forest conditions and to correspond with a concurrent ground plot sampling campaign. Airborne laser scanning data were processed into height-normalized point clouds and reprojected to a custom Lambert conformal conic projection to align with existing national satellite information products. More than 15 million 900 m2 lidar plots were generated from the 2023 transect dataset with point cloud metrics (i.e., area-based statistical summaries of the ALS point cloud) calculated for each 30 by 30 m cell. Presently, the 2023 lidar plots and their associated point cloud metrics are stored in openly available SQLite GeoPackages, with additional annual transect collections to be added when available. To accommodate a wide range of users and applications, both comprehensive and abridged versions of the metric databases, with 369 metrics and 40 metrics, respectively, are shared. The framework that led to the data shared here is portable to other areas with similar information needs. The data structure used was designed to enable updates with additional open access databases of ALS transects as data acquisition and processing are completed. This open-access dataset constitutes a vital resource for the scientific and operational forestry communities, offering detailed and scalable measures that bridge the gap between ground observations and wall-to-wall satellite-based inventories. These data will support the development of enhanced wildfire fuels maps, forest inventories, and carbon products. Acknowledgements The ALS data were acquired with funding from the Canadian Forest Service’s Northern Forest Mapping (NorthForM) program, which aims to enhance mapping of Canada's northern forests, identify wildfire hazards, and support community wildfire resilience and mitigation measures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.011 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".