Socially Intelligent Reinforcement Learning for Optimal Automated Vehicle Control in Traffic Scenarios
Bibliographic record
Abstract
In this paper, a novel approach is presented for modeling the interaction dynamics between an ego car and a bicycle in a traffic scenario using a hybrid reinforcement learning framework combined with a social value orientation (SVO) model. The proposed framework leverages the SARSA algorithm to learn the optimal policy for the ego vehicle while incorporating risk cost as the negative log-likelihood of collision. Additionally, a customized SVO model is introduced to capture the social preferences of the ego car and the bicycle, defining the SVO of each agent as a continuous variable between egoistic and cooperative orientations. Furthermore, a weight parameter is incorporated in the framework to regulate the influence of the SVO model on the learning process. We demonstrate the effectiveness of our approach through extensive simulations, showing that the ego car can balance between maximizing its reward and avoiding collisions while considering the social preferences of the agents. The obtained results are compared to other models in the literature, and it is shown that the proposed method contributes to the development of safe and efficient autonomous driving systems that interact with human-driven vehicles in a socially intelligent mannerNote to Practitioners—This proposed framework is motivated by the pressing challenge of navigation for autonomous cars in complex urban driving scenarios and mixed traffic situations. With the increasing prevalence of autonomous vehicles on roads, developing intelligent navigation systems that can effectively interact with other road users has become essential. Our novel framework addresses this need by leveraging the SARSA algorithm to learn the optimal policy for the ego vehicle while incorporating risk cost as the negative log-likelihood of collision. Additionally, a customized SVO model is introduced to capture the social preferences of the ego car and the bicycle, defining the SVO of each agent as a continuous variable between egoistic and cooperative orientations. This enables autonomous vehicles to make informed decisions and navigate safely and efficiently. Our framework can enormously help the field of autonomous vehicle navigation and contribute significantly to developing safe, human-centric, and reliable transportation systems. The versatility of our approach is evident in its potential to support a network of autonomous vehicles interacting with multiple road users, thereby enhancing scalability. By leveraging the power of machine learning, our solution provides a robust and adaptable approach that can handle the diverse and ever-changing conditions of urban driving scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".