Unsupervised Semantic Segmenting TLS Data of Individual Tree Based on Smoothness Constraint Using Open-Source Datasets
Bibliographic record
Abstract
Unsupervised segmentation of Terrestrial Laser Scanning (TLS) data into wood and leaf is the key for studying forest carbon storage, photosynthesis, canopy radiation. Further segmentation of wood data into trunk and larger branch (TLB), remaining branch (RB) is of great significance and challenge for dust retention, soil heavy metal enrichment. We proposed an unsupervised, automatic semantic segmentation method based on TLS data of individual tree. The method firstly performs initial segmentation based on plane fitting residuals and neighborhood normal angle, which can extract smooth and connected regions in point cloud. Then, the geometric features of segmented clusters are quantified to approximate RB or leaf features. Finally, the segmentation of TLB, RB, and leaf is realized by combining different clusters from bottom to top with geometric features and neighborhood relations. The segmentation performance of our method was evaluated with 104 tree samples from 23 tree species in two open-source datasets from Indonesia, Peru, Guyana and from Canada and Finland. The micro-average precision of our method is 93.61%. The micro-average recalls of TLB, RB, and leaf are 97.08%, 86.44%, and 89.62%. Compared with the well-known method of separating wood and leaf, our method has 33.56% higher sensitivity, 1.82% higher specificity, 20.52% higher precision, and 0.217 higher F1-score. Besides, we estimated the surface area and volume of TLB, the surface area and volume of RB based on the segmented data. The above parameters have good consistency compared to those calculated based on manually separated point clouds (Pearson correlation coefficient (PCC) of 0.55-0.93).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".