Synthesizing interpretable strategies for real-time planning in zero-sum games
Notice bibliographique
Résumé
Interpretable and explainable Artificial Intelligence (AI) is projected as one of the most important topics for the community in the next years. In addition to developing effective AI approaches that can help humans solving problems, it might be necessary to understand the reasons behind the decisions of such approaches to finally trust in their behavior. Search and learning-based algorithms represent the current state-of-the-art approaches for planning in zero-sum real-time games. The problem with those approaches is that usually the behavior of their resulting agents is not interpretable. On the other hand, hard-coded programs usually are not as effective as searchbased methods but have an important vantage; they can be more easily interpretable. In this thesis, we present a collection of works where we approach the problem of synthesizing effective interpretable scripts for planning in zero-sum real-time domains. First, we approach the problem of generating a set of scripts that can be used as an action abstraction to reduce search action spaces in zero-sum real-time strategy games. Namely, we present an evolutionary approach that can generate action abstractions that search-based algorithms can use for planning. Search-based systems that use action abstractions generated by our system outperformed the state-of-the-art search-based methods we use for experiments and won the 2018 RTS competition. We also present Gesy and LS2, two systems focused on synthesizing scripts that can plan by themselves in zero-sum real-time strategy games. Gesy is a system that uses a Genetic Programming (GP) approach to synthesize interpretable scripts. LS2 is a system that combines a novel method to reduce Domain-Specific Languages (DSLs), and a local-search algorithm that uses self play to synthesize interpretable scripts. The scripts Gesy and LS2 synthesize are competitive with complex search-based methods and scripts designed by professional programmers. We also show that the scripts synthesized by both systems can be used to discover possible optimizations that programmers could include in their implementations.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».