A Multi-Agent DRL-Based Dynamic Resource Allocation in O-RAN-Enabled TN-NTN Metaverse Services
Bibliographic record
Abstract
The integration of terrestrial and non-terrestrial networks (TN-NTN) with open radio access network (O-RAN) technology presents a significant advancement for facilitating scalable and immersive Metaverse services within 6G networks. Seamless virtual experiences necessitate highly reliable, low-latency communication, effective resource management, and adaptive decision-making to satisfy the varied and rigorous requirements of Metaverse applications, including gaming, healthcare, and autonomous systems. The inherent heterogeneity, dynamic nature, and substantial resource requirements of TN-NTN present significant challenges for effective resource allocation and optimizing quality of experience (QoE). Then, we formulate a multi-objective optimization problem for joint resource allocation and spectrum sharing in O-RAN-enabled TN-NTN Metaverse environments. This problem is inherently NP-hard due to the intricate coupling between continuous action spaces and discrete decision variables. Solving such a complex problem using traditional optimization approaches is complex. To overcome this, we transform the problem into a decentralized partially observable Markov decision process (Dec-POMDP) and address it using a hierarchical multi-agent deep reinforcement learning (MADRL) approach. This study presents a hierarchical multi-agent proximal policy optimization (MAPPO) framework, a new MADRL solution for dynamic resource allocation and spectrum sharing in O-RAN-enabled TN-NTN Metaverse environments. MAPPO facilitates collaborative learning among intelligent agents to optimize resource management strategies in a decentralized manner, considering essential metrics, including energy consumption, latency, and meta-distance. The proposed framework enhances resource utilization efficiency, minimizes latency, and improves the QoE for Metaverse users through the seamless allocation and management of resources. Comprehensive simulations show that MAPPO outperforms baseline methods, such as conventional reinforcement learning and centralized optimization approaches, achieving better energy efficiency, lower latency, and improved QoE. This demonstrates its effectiveness in adapting to dynamic 6G-enabled Metaverse requirements, enabling intelligent and scalable TN-NTN networks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".