MétaCan
Menu
Retour à la cohorte
Enregistrement W4388761312 · doi:10.17760/d20486919

Bayesian partially observable reinforcement learning

2023· dissertation· en· W4388761312 sur OpenAlexaff
Sammie Katt

Notice bibliographique

Revuenon disponible
Typedissertation
Langueen
DomaineComputer Science
ThématiqueReinforcement Learning in Robotics
Établissements canadiensScience North
Organismes subventionnairesnon disponible
Mots-clésComputer scienceReinforcement learningInferenceState (computer science)Bayesian inferenceArtificial intelligenceBayesian probability

Résumé

récupéré en direct d'OpenAlex

Autonomous agents are occupying more roles in our world than ever. They are present as AI in games, decide on which ads users see on the internet, and are even considered in more impactful environments such as finances, health, and perhaps traffic. As our expectations grow, their responsibilities become increasingly complex. To meet these demands, agents must reason over the current and potential future states of the environment. Reasoning over the state is not limited to understanding the latest observation but extends to inference over the hidden part of the environment. A natural solution is to maintain a probability distribution, or belief, over the state. Similarly, regarding future states, intelligent behavior considers the likelihood of multiple outcomes far into the future, also called planning. Both aspects of decision-making require full knowledge of the behavior, or dynamics, of the environment. These dynamics are used in planning to simulate future interactions. Similarly, imagined simulations of the system allow for reasoning about its current state. For example, a self-driving car may want to reason about the behavior of pedestrians for its decision-making (or indeed to estimate the likelihood of one being hidden from view!). Typically the dynamics are not fully known and, thus, these techniques are not directly applicable in practice. When the dynamics are not available, then the best thing we can do is to provide a prior over the unknown quantities instead. For example, even if there is no fully accurate model of pedestrians, we may provide a distribution over possible behaviors. This approach, called Bayesian reinforcement learning, allows us to provide the agent with expert knowledge. With this setup, it is now possible to update our understanding of the unknown quantities of the system when new data becomes available. The result is an inference problem where we maintain a belief over both the state and the dynamics of the environment. This Bayesian reinforcement learning approach was best formalized in the literature with the Bayes-adaptive models. Unfortunately, although elegant in theory, the approach had some limitations and saw little adoption in practice. In particular, no planner had been developed to allow for efficient action selection. Second, the type of prior knowledge that the proposed models were able to capture restricted their usage to small problems. Variants with more complex representations required strong priors and the structure of the dynamics would be assumed known. This thesis advances Bayesian partially observable reinforcement learning to non-trivial domains with several contributions. In particular, it discusses a holistic definition of the Bayesian inference problem, improved planning algorithms, and scalable model representations plus approaches for their approximation. First, we visualize the inference problem over state and dynamics with a graphical model and exploit its structure to define a novel posterior derivation. Not only does this derivation unify previous work into a single recipe, but also opens the door to other (machine learning) approaches. The resulting framework is formalized as the general Bayes-adaptive Markov decision process'' (GBA-POMDP). The second contribution is a family of efficient planning algorithms that were the first technique that made it possible to do decision-making in the Bayes-adaptive models. These planners are specializations of Monte-Carlo tree search, which are sample-based methods that approximate the value of future actions through simulations. In particular, we combine additional sampling approximations with inference simplifications to make reasoning over the complex GBA-POMDP possible. Third, we discuss two sophisticated models that move away from the tabular representations that were assumed in previous work. We show how a Bayes-net representation allows us to model and learn structure in the dynamics of the system. This includes the usage of an intricate and targeted sampling scheme to avoid the collapse of the posterior approximation. Lastly, we show the practical use of the GBA-POMDP withBayes-Adaptive Deep Dropout Reinforcement learning (BADDr)'', which employs neural networks to model the dynamics of the environment. BADDr, with the expressiveness and scalability of neural networks, showcases how Bayesian reinforcement learning can be both principled and practical.--Author's abstract

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,011
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,011
Score d'incertitude au seuil0,021

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,011
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0010,001
Études des sciences et des technologies0,0010,002
Communication savante0,0010,002
Science ouverte0,0020,002
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,0060,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,025
Tête enseignante GPT0,273
Écart entre enseignants0,247 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même sujetReinforcement Learning in RoboticsTravaux en français237 207