Deep Reinforcement Learning based Model-Free Wide-Area Damping Control under Uncertain Delays
Bibliographic record
Abstract
Although inter-area signals measured by PMU provides more robust control to mitigate the oscillations occurring at lower frequencies, controlling techniques relying on wide-area signals encounter considerable obstacles posed by transmission delays over longer communication channels stretching from PMUs to wide-area damping control (WADC) systems. Managing fluctuating transmission delays is crucial, particularly within expansive monitoring and control systems. Without any knowledge of the real-time transmission delays at every sample time, this paper presents pioneering method employing Deep Reinforcement Learning (DRL) for WADC, devoid of traditional modeling, to effectively mitigate inter-area low-frequency oscillations. The state set includes both cyber layer information and physical layer power system performance. The optimal control policy is designed by deep determinisitic policy gradient (DDPG) to maximize the reward function on every sample time. The reward function is designed based on the working features of the generators for enabling timely damping oscillations. The IEEE 10-Generator 39-Bus system serves as the focal point for comparative analyses conducted in this study using the proposed model-free DDPG based WADC with the proposed state set and with conventional state set. The results show that based on the proposed state set, DRL agent can get enough information of the environment. The trained optimal control policy can be continuously updated by DDPG agent with a maximum reward value. With the proposed state set, the ddpg based WADC system exemplifies the rapid suppression of inter-area low-frequency oscillations and the dependable stabilization of the power system, even when faced with uncertain time delays.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".