A Proximal Policy Optimization-Based Controller for Enhanced Power Sharing in Microgrids
Bibliographic record
Abstract
This paper introduces a Proximal Policy Optimization (PPO)-based virtual impedance (VI) controller to enhance both power sharing and system response under disturbances in inverter-interfaced microgrids. Traditional droop control methods often face challenges due to variations in feeder impedance, which degrade performance. The proposed controller continuously updates its policy based on changes in the operating environment. The control problem is modeled as a Markov Decision Process (MDP), in which the state and action spaces are explicitly defined, and a carefully designed reward function, satisfying system criteria and constraints, guides the learning process toward achieving the desired transient and steady-state performance. By leveraging PPO, the controller improves upon traditional methods by reducing the need for manual tuning and offering better adaptability to varying operating conditions. The performance of the proposed controller is evaluated in both islanded and grid-connected modes, using batteries with capacities of 1 MW, 125 kW, and 100 kW. The results demonstrate that the PPO-based VI controller improves power-sharing accuracy and provides better response to disturbances across different scenarios compared to the conventional controller. To validate the performance of the proposed method, an assessment is conducted on the system frequency using key metrics, including Root Mean Square Error (RMSE), Integral of Absolute Error (IAE), Integral of Squared Error (ISE), and Integral of Time-Weighted Squared Error (ITSE). The PPO controller consistently achieves the lowest errors across all scenarios compared to the conventional controller, with the IAE reduced by 27% in islanded mode and 36% in grid-connected mode.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".