Optimizing Robotic Arm Control Using Deep Deterministic Policy Gradient: An Exploration of Hyperparameter Tuning
Bibliographic record
Abstract
Robotic arms are essential in a wide range of applications, from industrial automation to medical surgeries, where both accuracy and adaptability are critical. Traditional path-planning methods for robotic arms, such as Rapidly-exploring Random Trees (RRT) and Probabilistic Roadmaps (PRM), often suffer from limitations in dynamic environments. Reinforcement Learning (RL) presents a promising alternative for optimizing robotic arm control by enabling adaptive learning through trial and error. This study focuses on the application of Deep Deterministic Policy Gradient (DDPG), a popular RL algorithm, to control a simulated robotic arm following a mouse pointer. The study investigates the impact of three key hyperparameters—learning rate, batch size, and memory capacity—on the performance of the DDPG model. This paper systematically tested multiple values for each parameter and evaluated the model's success rate and average time per goal. Results showed that the optimal combination of parameters was a learning rate of 0.001, a batch size of 50, and a memory capacity of 30,000, yielding a success rate of 76.00% and an average time per goal of 0.07 seconds. These results emphasize the significance of fine-tuning hyperparameters to achieve optimal performance in robotic control tasks. Future work will focus on exploring adaptive hyperparameter tuning strategies and applying these methods to more complex and dynamic robotic environments to further enhance performance and adaptability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".