A Comparative Study Between Soft Actor-Critic (SAC) and Deep Deterministic Policy Gradient (DDPG) Algorithms for Solar PV MPPT Control Under Partial Shading Conditions
Notice bibliographique
Résumé
The use of photovoltaic (PV) arrays in smart grid systems is growing due to the increasing energy demand and greenhouse gas emissions. However, due to the intermittent nature of PV arrays, the Maximum Power Point Tracking (MPPT) algorithm is typically employed to optimize the system’s energy production. In the past, the conventional perturb and observe (P&O) method was proposed for solar PV MPPT control. While the P&O method can estimate the PV maximum power under uniform irradiation, it often exhibits sluggish tracking and unstable steady-state oscillations and fails to track the global maximum power point (GMPP) under partial shading conditions (PSCs). These problems have been addressed using deep reinforcement learning (RL) algorithms, such as the deep deterministic policy gradient (DDPG) algorithm. However, due to the DDPG’s intrinsic drawbacks, such as unstable training, Q-value overestimation, brittle convergence, and hyperparameter sensitivity, it often produces steady-state power oscillations near the GMPP under PSCs, resulting in power loss. This paper presents a soft actor-critic (SAC) algorithm, for solving solar PV MPPT control problems under PSCs. Unlike DDPG, which utilizes only one Q-network in the critic, SAC utilizes two Q-networks in the critic and maximum entropy policy in the reward function, which guarantees its training stability and improves its exploration and robustness in the presence of “estimation and model errors”. Despite its potential, the SAC-based MPPT approach has not been extensively explored or compared with DDPG to determine the superior method for PV MPPT control. This paper provides an adequate comparison between the performances of DDPG and SAC, including their optimal hyperparameter configurations, for PV MPPT control. To solve the MPPT control problem, the mathematical model of the boost converter and the solar PV system were developed. Then, a Markov Decision Process model was formulated, which represents the PV system’s behavior. For completeness in the comparison, the conventional P&O algorithm was also included. Simulation results show that SAC and DDPG algorithms achieved superior performance compared to the P&O method under PSCs, and constant and varying irradiance levels. It is shown that the SAC algorithm provides superior performance in achieving high tracking efficiency and zero power oscillations near the PV MPP and GMPP compared to the DDPG method.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,005 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».