Implicit Representation-based Volumetric Video Streaming for Photorealistic Full-scene Experience
Bibliographic record
Abstract
The widespread integration of the Internet of Things with sensors like depth-of-field cameras, LiDAR scanners, and eye-tracking infrared sensors, in head-mounted devices, has ushered in a new era of immersive digital experiences. Full-scene volumetric video (VV), a key innovation in this integration, provides a deeply immersive experience by capturing the richness and detail of the 3D world. However, its massive data volume presents significant streaming challenges. While 3D tile-based viewport approaches have been proposed, they struggle to full-scene VV given the small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration. In this work, inspired by the advancements of implicit neural radiance field (NeRF), we present \({\mathsf{V}^{2}\mathsf{NeRF}}\) , a novel full-scene VV streaming system featured by layered representation. It harmonizes the NeRF with explicit point clouds to represent the static background and dynamic foreground, thereby avoiding large data transfers and achieving photorealistic content representation. To tackle the issues of intensive computation requirements and multiscale adaptation scheduling within \({\mathsf{V}^{2}\mathsf{NeRF}}\) system, we propose a lightweight non-visible background removal method and a two-stage decoupled architecture. In addition, an efficient buffer-aware simulated annealing algorithm is developed, alongside the utilization of a perceptually learned metric, to enhance user experience. We further discuss the concerns about practical development and deployment. Extensive prototype evaluations demonstrate \({\mathsf{V}^{2}\mathsf{NeRF}}\) ’s superior streaming and viewing performance on a wide variety of networks, viewing motions, and scenes. For instance, compared to state-of-the-art approaches, it achieves a 24% increment in perceptual quality, an 83% reduction in rebuffering time, and a 54% enhancement in user experience on average.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".