Bibliographic record
Abstract
Abstract. Light Detection and Ranging (LiDAR) systems are used intensively in terrain surface modelling based on the range data determined by the LiDAR sensors. LiDAR sensors record the distance between the sensor and the targets (range data) with a capability to record the strength of the backscatter energy reflected from the targets (intensity data). The LiDAR sensors use the near-infrared spectrum range which has high separability in the reflected energy from different targets. This characteristic is investigated to implement the LiDAR intensity data in land-cover classification. The goal of this paper is to investigate and evaluates the use of LiDAR data only (range and intensity data) to extract land cover information. Different bands generated from the LiDAR data (Normal Heights, Intensity Texture, Surfaces Slopes, and PCA) are combined with the original data to study the influence of including these layers on the classification accuracy. The Maximum likelihood classifier is used to conduct the classification process for the LiDAR Data as one of the best classification techniques from literature. A study area covering an urban district in Burnaby, British Colombia, Canada, is selected to test the different band combinations to extract four information classes: buildings, roads and parking areas, trees, and low vegetation (grass) areas. The results show that an overall accuracy of more than 70% can be achieved using the intensity data, and other auxiliary data generated from the range and intensity data. Bands of the Principle Component Analysis (PCA) are also created from the LiDAR original and auxiliary data. Similar overall accuracy of the results can be achieved using the four bands extracted from the Principal Component Analysis (PCA).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".