Deep Reinforcement Learning Agents for Decision Making for Gameplay
Bibliographic record
Abstract
Robots are becoming more integrated into society as they become more advanced, and the programming behind them needs to continue to progress in order for the robots to be utilized to their fullest potential. Artificial Intelligence (AI) is one of the most versatile and quickly growing areas of robotic control, and has been used for a variety of different robots and tasks. One potential use of robotics and AI is in that of childhood development. Cooperative play has been shown to be a crucial part of childhood development, and for children with developmental disabilities, playing with other children may be difficult and frustrating, leading them to miss out on this important milestone. Cooperative play with robots has been shown to have positive educational and therapeutic effects on children with developmental disabilities, and so robots can be used as substitute players for children who have troubles playing with other children. To achieve this, AI algorithms must be developed that can make appropriate decisions or moves for a given game, to such an extent that the children would choose to play with the robot instead of alone. In this paper two AI agents are developed to play Menara, a cooperative tower building game. The two agents include a pillar placement agent and a tile placement agent. They implement algorithms including the method for selecting the pillars to have available to the agent during gameplay, and how many pillars the agent plans to place in a single turn. The tile placement agent was able to successfully balance a tile 62% of the time, while the pillar placement agent was able to succeed 88% of the time on the test dataset.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".