Accelerating Reinforcement Learning via Predictive Policy Transfer in 6G RAN Slicing
Bibliographic record
Abstract
Reinforcement Learning (RL) algorithms have recently been proposed to solve dynamic radio resource management (RRM) problems in beyond 5G networks. However, RL-based solutions are still not widely adopted in commercial cellular networks. One of the primary reasons for this is the slow convergence of RL agents when they are deployed in a live network and when the network’s context changes significantly. Concurrently, the open radio access network (O-RAN) paradigm promises to give mobile network operators (MNOs) more control over their networks, furthering the need for intelligent and RL-based network management. O-RAN’s standardized interfaces will allow MNOs to make real-time custom changes to intelligently control various RRM functionalities. We consider a RAN slicing scenario in which MNOs can modify the weights of the RL reward function. This enables MNOs to change the priorities of fulfilling the service level agreements of the slices. However, this results in a practical challenge since the RL agent needs to adapt promptly to the changes made by the MNO. This challenge is addressed in this paper, where we first present and discuss the results from an exhaustive experiment to examine the efficiency of using transfer learning (TL) to accelerate the convergence of RL-based RAN slicing in the considered scenario. We then propose a novelpredictiveapproach to enhance the TL-based acceleration by selecting the best-saved policy for reuse. By adopting the proposed policy transfer approach, RL agents are able to converge up to 14000 learning steps faster than their non-accelerated counterparts. The proposed machine learning (ML)-basedpredictiveapproach also shows up to a 96.5% accuracy in selecting the best expert policy to reuse for acceleration.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".