A New Model-Free Space Vector Modulation Technique for Multilevel Inverters Based On Deep Reinforcement Learning
Bibliographic record
Abstract
Multilevel inverters provide various advantages over conventional inverters. However, control and switching these types of inverters is a challenge for an increased number of levels that necessitate a higher number of semiconductor devices. Space Vector Modulation (SVM) switching method is a popular one that provides several advantages over conventional Sinusoidal Pulse Width Modulation techniques. Advantages such as higher output voltage, lower switching loss, lower dv/dt, etc. Despite its advantages, the SVM switching method requires complex time-interval equations. This complexity increases exponentially by increasing the number of levels. In this paper, a new SVM method based on Deep Reinforcement Learning (DRL) is proposed. The latter is based on a Reinforcement Learning Agent (RL-Agent) that tries all the "allowed" switching states of an inverter regarding the reference signal in each time step. Then based on the accuracy of the output, the agent receives a reward or punishment (negative reward). Over several episodes of trial and error, the agent learns how to perform the switching to receive the maximum long-term reward. Since the agent is also rewarded for keeping the voltage of the DC-link capacitors in a given range, this method is capable of voltage balancing. This method is implemented on a three-phase three-level Neutral Point Clamped (NPC) inverter for performance evaluation. The RL-agent is trained in MATLAB while the inverter model is simulated in the Simulink environment. The simulation results demonstrate the decent performance of this method.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".