Evaluation of LiDAR-Derived Features Relevance and Training Data Minimization for 3D Point Cloud Classification
Bibliographic record
Abstract
Terrestrial laser scanning (TLS) is a leading technology in data acquisition for building information modeling (BIM) applications due to its rapid, direct, and accurate scanning of different objects with high point density. Three-dimensional point cloud classification is essential step for Scan-to-BIM applications that requires high accuracy classification methods, running at reasonable processing time. The classification process is divided into three main steps: neighborhood definition, LiDAR-derived features extraction, and machine learning algorithms being applied to label each LiDAR point. However, the extraction of LiDAR-derived features and training data are time consuming. This research aims to minimize the training data, assess the relevance of sixteen LiDAR-derived geometric features, and select the most contributing features to the classification process. A pointwise classification method based on random forests is applied on the 3D point cloud of a university campus building collected by a TLS system. The results demonstrated that the normalized height feature, which represented the absolute height above ground, was the most significant feature in the classification process with overall accuracy more than 99%. The training data were minimized to about 10% of the whole dataset with achieving the same level of accuracy. The findings of this paper open doors for BIM-related applications such as city digital twins, operation and maintenance of existing structures, and structural health monitoring.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".