A novel quality control procedure for the evaluation of laser scanning data segmentation
Bibliographic record
Abstract
Abstract. Over the past few years, laser scanning systems have been acknowledged as the leading tools for the collection of high density 3D point cloud over physical surfaces for many different applications. However, no interpretation and scene classification is performed during the acquisition of these datasets. Consequently, the collected data must be processed to extract the required information. The segmentation procedure is usually considered as the fundamental step in information extraction from laser scanning data. So far, various approaches have been developed for the segmentation of 3D laser scanning data. However, none of them is exempted from possible anomalies due to disregarding the internal characteristics of laser scanning data, improper selection of the segmentation thresholds, or other problems during the segmentation procedure. Therefore, quality control procedures are required to evaluate the segmentation outcome and report the frequency of instances of expected problems. A few quality control techniques have been proposed for the evaluation of laser scanning segmentation. These approaches usually require reference data and user intervention for the assessment of segmentation results. In order to resolve these problems, a new quality control procedure is introduced in this paper. This procedure makes hypotheses regarding potential problems that might take place in the segmentation process, detects instances of such problems, quantifies the frequency of these problems, and suggests possible actions to remedy them. The feasibility of the proposed approach is verified through quantitative evaluation of planar and linear/cylindrical segmentation outcome from two recently-developed parameter-domain and spatial-domain segmentation techniques.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".