Notice bibliographique
Résumé
Autonomous agents are occupying more roles in our world than ever. They are present as AI in games, decide on which ads users see on the internet, and are even considered in more impactful environments such as finances, health, and perhaps traffic. As our expectations grow, their responsibilities become increasingly complex. To meet these demands, agents must reason over the current and potential future states of the environment. Reasoning over the state is not limited to understanding the latest observation but extends to inference over the hidden part of the environment. A natural solution is to maintain a probability distribution, or belief, over the state. Similarly, regarding future states, intelligent behavior considers the likelihood of multiple outcomes far into the future, also called planning. Both aspects of decision-making require full knowledge of the behavior, or dynamics, of the environment. These dynamics are used in planning to simulate future interactions. Similarly, imagined simulations of the system allow for reasoning about its current state. For example, a self-driving car may want to reason about the behavior of pedestrians for its decision-making (or indeed to estimate the likelihood of one being hidden from view!). Typically the dynamics are not fully known and, thus, these techniques are not directly applicable in practice. When the dynamics are not available, then the best thing we can do is to provide a prior over the unknown quantities instead. For example, even if there is no fully accurate model of pedestrians, we may provide a distribution over possible behaviors. This approach, called Bayesian reinforcement learning, allows us to provide the agent with expert knowledge. With this setup, it is now possible to update our understanding of the unknown quantities of the system when new data becomes available. The result is an inference problem where we maintain a belief over both the state and the dynamics of the environment. This Bayesian reinforcement learning approach was best formalized in the literature with the Bayes-adaptive models. Unfortunately, although elegant in theory, the approach had some limitations and saw little adoption in practice. In particular, no planner had been developed to allow for efficient action selection. Second, the type of prior knowledge that the proposed models were able to capture restricted their usage to small problems. Variants with more complex representations required strong priors and the structure of the dynamics would be assumed known. This thesis advances Bayesian partially observable reinforcement learning to non-trivial domains with several contributions. In particular, it discusses a holistic definition of the Bayesian inference problem, improved planning algorithms, and scalable model representations plus approaches for their approximation. First, we visualize the inference problem over state and dynamics with a graphical model and exploit its structure to define a novel posterior derivation. Not only does this derivation unify previous work into a single recipe, but also opens the door to other (machine learning) approaches. The resulting framework is formalized as the general Bayes-adaptive Markov decision process'' (GBA-POMDP). The second contribution is a family of efficient planning algorithms that were the first technique that made it possible to do decision-making in the Bayes-adaptive models. These planners are specializations of Monte-Carlo tree search, which are sample-based methods that approximate the value of future actions through simulations. In particular, we combine additional sampling approximations with inference simplifications to make reasoning over the complex GBA-POMDP possible. Third, we discuss two sophisticated models that move away from the tabular representations that were assumed in previous work. We show how a Bayes-net representation allows us to model and learn structure in the dynamics of the system. This includes the usage of an intricate and targeted sampling scheme to avoid the collapse of the posterior approximation. Lastly, we show the practical use of the GBA-POMDP withBayes-Adaptive Deep Dropout Reinforcement learning (BADDr)'', which employs neural networks to model the dynamics of the environment. BADDr, with the expressiveness and scalability of neural networks, showcases how Bayesian reinforcement learning can be both principled and practical.--Author's abstract
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,002 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».