LKG-Net: Local-Keypoint-Global Feature Fusion Network for Sparse Point Cloud Semantic Segmentation
Bibliographic record
Abstract
Semantic segmentation of sparse point clouds presents a significant challenge in 3D computer vision. Sparsity primarily arises from two sources: suboptimal data quality due to limitations in point cloud acquisition devices or environmental factors, and weakly supervised learning strategies employed to reduce annotation costs. Existing methods often struggle to simultaneously achieve accurate perception of fine-grained local geometric structures and effective understanding of global scene context when processing sparse point clouds. To address this challenge, we propose LKG-Net, a novel end-to-end segmentation network that systematically enhances feature learning and fusion capabilities through three carefully designed modules. The Adaptive Hierarchical Efficient MLP Module (AHEM) enables deep extraction of robust local features through hierarchical feature refinement and adaptive pooling strategies. The Keypoint-Driven Local-to-Global Spatial Feature Module (KGSF) employs an efficient keypoint-based attention mechanism to capture global contextual information while significantly reducing computational complexity. The Local-Global Fusion Module (LGF) dynamically and adaptively merges multi-scale features based on data characteristics, ensuring optimal feature integration across different spatial scales. Comprehensive experiments on the S3DIS and Toronto-3D benchmark datasets demonstrate the effectiveness of our method. Specifically, LKGNet establishes a new state of the art on the S3DIS Area 5 test set under an extremely sparse (0.1 %) supervision setting, achieving a mean Intersection over Union (mIoU) of <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{7 0. 4 \%}$</tex> and surpassing strong baselines by 2.7 %. On Toronto-3D, LKG-Net further demonstrates robust generalization with superior performance on complex outdoor scenes. These results conclusively establish LKG-Net's exceptional capability for precise segmentation under sparse point cloud conditions, providing a robust and generalizable solution for sparse supervised 3D scene understanding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".