LiDAR few-shot domain adaptation via integrated CycleGAN and 3D object detector with joint learning delay
Bibliographic record
Abstract
he success of supervised LiDAR perception methods relies on the availability of large sets of labeled point cloud data, for which the labeling process is costly and time consuming. Given unpaired LiDAR datasets of similar sizes from two domains, with one (source) containing task-specific labels e.g. 3D bounding boxes for all frames, but only a small percentage of frames being labeled in the other (target) domain, it is challenging to train a model that generalizes well on validation data from the target domain. In this paper we propose a novel LiDAR few-shot domain adaptation architecture and training strategy to address this challenge. Our method is based on adapting a task-specific network (3D object detector) to work within the CycleGAN framework modified to operate with LiDAR features, and on the joint end-to-end training of generators, discriminators, and task-specific layers. To overcome nonconvergence issues we propose a training strategy that introduces a mechanism to delay the joint learning between the generators/discriminators and the task-specific network by allowing them to start learning independently, while slowly introducing joint learning as they converge, hence avoiding instability during the early stages of the training. Our proposed integrated architecture enables a direct way to evaluate the performance of the model instead of feeding pre-computed generated data into a separate pretrained model. We include an experimental section where we evaluate our proposed architecture on the publicly available KITTI and Nuscenes datasets, as well as on our own labeled dataset. We present useful mean average precision plots that illustrate the benefits of our domain adaptation architecture as a function of number of labeled target domain frames.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".