Classification of Multispectral Airborne LiDAR Data Using Geometric and Radiometric Information
Bibliographic record
Abstract
Classification of airborne light detection and ranging (LiDAR) point cloud is still challenging due to the irregular point cloud distribution, relatively low point density, and the complex urban scenes being observed. The availability of multispectral LiDAR systems allows for acquiring data at different wavelengths with a variety of spectral information from land objects. In this research, a rule-based point classification method of three levels for multispectral airborne LiDAR data covering urban areas is presented. The first level includes ground filtering, which attempts to distinguish aboveground from ground points. The second level aims to divide the aboveground and ground points into buildings, trees, roads, or grass using three spectral indices, namely normalized difference feature indices (NDFIs). A multivariate Gaussian decomposition is then used to divide the NDFIs’ histograms into the aforementioned four classes. The third level aims to label more classes based on their spectral information such as power lines, types of trees, and swimming pools. Two data subsets were tested, which represent different complexity of urban scenes in Oshawa, Ontario, Canada. It is shown that the proposed method achieved an overall accuracy up to 93%, which is increased to over 98% by considering the spatial coherence of the point cloud.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".