Two-Time-Scale Hierarchical Sequential Multiagent DRL Framework With Compute Reuse for Energy and Delay Minimization in UAV-Based Edge
Bibliographic record
Abstract
Unmanned aerial vehicles (UAV)-based edge computing (EC) has been a core enabler for task offloading in the last decade with a major focus on UAV energy and task processing delay minimization, especially for delay-sensitive applications that require high availability of task offloading services throughout the day, such as vehicular networks. In this work, we tackle the problem of task offloading for internet of things (IoT) devices in UAV-based EC networks by introducing compute reuse for result sharing. The goal is to minimize the energy consumption of the edge-computing UAVs and the task processing delay. By considering the original problem split into two subproblems on two different time scales: the small time scale handles the forwarding decisions and the CPU operating frequency at each UAV, while the large time scale handles the wake-sleep status of each UAV. Most existing work in this area ignores the discussion of result reuse in a UAV edge setup and considers a single agent on each time scale without sequential decision-making. Additionally, the literature concerning sending UAVs to sleep for energy savings under edge computing framework or the integration of both sequentiality and hierarchical deep reinforcement learning (DRL) under a multi-time scale framework remains unexplored in the literature. An algorithm based on a hierarchical sequential multi-agent deep reinforcement learning (HSMADRL) framework is developed. The hierarchical DRL explores patterns between the two time scales; and the sequential and multi-agent designs address the curse of dimensionality and decision-making coordination, respectively. Simulation results show that our algorithm can quickly converge and outperform baselines, while the energy computation efficiency (ECE) is maximized.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".