A Reinforcement Learning Congestion Control Algorithm for Smart Grid Networks
Bibliographic record
Abstract
Modern electrical systems are evolving with data communication networks, ushering in upgraded electrical infrastructures and enabling bidirectional communication between utility grids and consumers. The selection of communication technologies is crucial, where wireless communications have emerged as one of the main benefactor technologies due to its cost-effectiveness, scalability, and ease of deployment. Ensuring the streamlined transmission of diverse applications between residential users and utility control centers is crucial for the effective data delivery in smart grids. This paper proposes a congestion control mechanism tailored to smart grid applications using unreliable transport protocols such as UDP, which, unlike TCP, lacks inherent congestion control, presenting a significant challenge to the performance of the network. In this article, we have exploited a reinforcement learning (RL) algorithm and a deep <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Q</i> -neural network (DQN), to manage congestion control under a UDP environment. In this particular instance, the DQN model learns on its own from interactions with the environment without the need to generate a dataset, rendering it apt for intricate and dynamic scenarios within smart grid communications. Our evaluation covers two scenarios: (i) a grid-like configuration, and (ii) urban scenarios considering the deployment of smart meters in the cities of Montreal, Berlin and Beijing. These evaluations provide a comprehensive examination of the proposed DQN-based congestion control approach under different conditions, showing its effectiveness and adaptability. Conducting a comprehensive performance assessment in both scenarios leads to improvements in metrics such as packet delivery ratio, network throughput, fairness between different traffic sources, packet network transit time, and QoS provision.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".