An Efficient Neural Network Architecture and Training Protocol for 3D Point Cloud Classification
Bibliographic record
Abstract
The point cloud is a set of data points in a 3D coordinate system with an irregular data format. As a result, they are needed to be transformed into a collection of images before being fed into models. This unnecessarily increases the volume of the data and increases complexities. The existing literature on point cloud uses a fixed number of points sampled from the whole point cloud as the input. However, with large point cloud data, it is important to consider more points as input to have a better understanding of the scene. The computational expense increases if the input number of points increases for existing networks. Our research contributes to the existing point cloud classification literature in two directions. First, we develop a training protocol for improved point cloud training accuracy on top of the existing PointNet \\cite{qi2017pointnet} architecture over the ModelNet10 dataset. A few variations of encoder models have been proposed in this regard. Also, an extensive hyperparameter study and ablation study are done. These experiments achieve a 6.10\\% improvement over the baseline model. After that, we propose DualNet, a novel 3D point cloud network that resolves the trade-off between the number of input points and the computational expense of 3D data. The DualNet consists of two branches: DensetNet and SparseNet. The SparseNet is a comparatively large network in terms of number of parameters, that samples a small number of points from the whole point cloud. Whereas the DenseNet is a lightweight network that takes a large number of points as input. SparseNet is composed of more number of channels than DenseNet making it more computationally expensive than DenseNet. While the accuracy of the model shows good improvement when the number of points increases, the overall computational cost of DenseNet does not increase much in such settings. DualNet shows 0.81\\% and 0.45\\% increase in the SOTA results on ModelNet40 and ScanObjectNN respectively. In respect of computational complexity, our model takes about 40\\% less time compared to SOTA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".