MétaCan
Menu
Retour à la cohorte
Enregistrement W4388761312 · doi:10.17760/d20486919

Bayesian partially observable reinforcement learning

2023· dissertation· en· W4388761312 sur OpenAlexaff
Sammie Katt

Notice bibliographique

Revuenon disponible
Typedissertation
Langueen
DomaineComputer Science
ThématiqueReinforcement Learning in Robotics
Établissements canadiensScience North
Organismes subventionnairesnon disponible
Mots-clésComputer scienceReinforcement learningInferenceState (computer science)Bayesian inferenceArtificial intelligenceBayesian probability

Résumé

récupéré en direct d'OpenAlex

Autonomous agents are occupying more roles in our world than ever. They are present as AI in games, decide on which ads users see on the internet, and are even considered in more impactful environments such as finances, health, and perhaps traffic. As our expectations grow, their responsibilities become increasingly complex. To meet these demands, agents must reason over the current and potential future states of the environment. Reasoning over the state is not limited to understanding the latest observation but extends to inference over the hidden part of the environment. A natural solution is to maintain a probability distribution, or belief, over the state. Similarly, regarding future states, intelligent behavior considers the likelihood of multiple outcomes far into the future, also called planning. Both aspects of decision-making require full knowledge of the behavior, or dynamics, of the environment. These dynamics are used in planning to simulate future interactions. Similarly, imagined simulations of the system allow for reasoning about its current state. For example, a self-driving car may want to reason about the behavior of pedestrians for its decision-making (or indeed to estimate the likelihood of one being hidden from view!). Typically the dynamics are not fully known and, thus, these techniques are not directly applicable in practice. When the dynamics are not available, then the best thing we can do is to provide a prior over the unknown quantities instead. For example, even if there is no fully accurate model of pedestrians, we may provide a distribution over possible behaviors. This approach, called Bayesian reinforcement learning, allows us to provide the agent with expert knowledge. With this setup, it is now possible to update our understanding of the unknown quantities of the system when new data becomes available. The result is an inference problem where we maintain a belief over both the state and the dynamics of the environment. This Bayesian reinforcement learning approach was best formalized in the literature with the Bayes-adaptive models. Unfortunately, although elegant in theory, the approach had some limitations and saw little adoption in practice. In particular, no planner had been developed to allow for efficient action selection. Second, the type of prior knowledge that the proposed models were able to capture restricted their usage to small problems. Variants with more complex representations required strong priors and the structure of the dynamics would be assumed known. This thesis advances Bayesian partially observable reinforcement learning to non-trivial domains with several contributions. In particular, it discusses a holistic definition of the Bayesian inference problem, improved planning algorithms, and scalable model representations plus approaches for their approximation. First, we visualize the inference problem over state and dynamics with a graphical model and exploit its structure to define a novel posterior derivation. Not only does this derivation unify previous work into a single recipe, but also opens the door to other (machine learning) approaches. The resulting framework is formalized as the general Bayes-adaptive Markov decision process'' (GBA-POMDP). The second contribution is a family of efficient planning algorithms that were the first technique that made it possible to do decision-making in the Bayes-adaptive models. These planners are specializations of Monte-Carlo tree search, which are sample-based methods that approximate the value of future actions through simulations. In particular, we combine additional sampling approximations with inference simplifications to make reasoning over the complex GBA-POMDP possible. Third, we discuss two sophisticated models that move away from the tabular representations that were assumed in previous work. We show how a Bayes-net representation allows us to model and learn structure in the dynamics of the system. This includes the usage of an intricate and targeted sampling scheme to avoid the collapse of the posterior approximation. Lastly, we show the practical use of the GBA-POMDP withBayes-Adaptive Deep Dropout Reinforcement learning (BADDr)'', which employs neural networks to model the dynamics of the environment. BADDr, with the expressiveness and scalability of neural networks, showcases how Bayesian reinforcement learning can be both principled and practical.--Author's abstract

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMéta-épidémiologie (sens strict), Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Autre · Signal consensuel: aucune
Score de désaccord entre enseignants0,848
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0010,000
Science ouverte0,0020,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,002

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,025
Tête enseignante GPT0,273
Écart entre enseignants0,247 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreAutre

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même sujetReinforcement Learning in RoboticsTravaux en français237 207