Bibliographic record
Abstract
Autonomous agents are occupying more roles in our world than ever. They are present as AI in games, decide on which ads users see on the internet, and are even considered in more impactful environments such as finances, health, and perhaps traffic. As our expectations grow, their responsibilities become increasingly complex. To meet these demands, agents must reason over the current and potential future states of the environment. Reasoning over the state is not limited to understanding the latest observation but extends to inference over the hidden part of the environment. A natural solution is to maintain a probability distribution, or belief, over the state. Similarly, regarding future states, intelligent behavior considers the likelihood of multiple outcomes far into the future, also called planning. Both aspects of decision-making require full knowledge of the behavior, or dynamics, of the environment. These dynamics are used in planning to simulate future interactions. Similarly, imagined simulations of the system allow for reasoning about its current state. For example, a self-driving car may want to reason about the behavior of pedestrians for its decision-making (or indeed to estimate the likelihood of one being hidden from view!). Typically the dynamics are not fully known and, thus, these techniques are not directly applicable in practice. When the dynamics are not available, then the best thing we can do is to provide a prior over the unknown quantities instead. For example, even if there is no fully accurate model of pedestrians, we may provide a distribution over possible behaviors. This approach, called Bayesian reinforcement learning, allows us to provide the agent with expert knowledge. With this setup, it is now possible to update our understanding of the unknown quantities of the system when new data becomes available. The result is an inference problem where we maintain a belief over both the state and the dynamics of the environment. This Bayesian reinforcement learning approach was best formalized in the literature with the Bayes-adaptive models. Unfortunately, although elegant in theory, the approach had some limitations and saw little adoption in practice. In particular, no planner had been developed to allow for efficient action selection. Second, the type of prior knowledge that the proposed models were able to capture restricted their usage to small problems. Variants with more complex representations required strong priors and the structure of the dynamics would be assumed known. This thesis advances Bayesian partially observable reinforcement learning to non-trivial domains with several contributions. In particular, it discusses a holistic definition of the Bayesian inference problem, improved planning algorithms, and scalable model representations plus approaches for their approximation. First, we visualize the inference problem over state and dynamics with a graphical model and exploit its structure to define a novel posterior derivation. Not only does this derivation unify previous work into a single recipe, but also opens the door to other (machine learning) approaches. The resulting framework is formalized as the general Bayes-adaptive Markov decision process'' (GBA-POMDP). The second contribution is a family of efficient planning algorithms that were the first technique that made it possible to do decision-making in the Bayes-adaptive models. These planners are specializations of Monte-Carlo tree search, which are sample-based methods that approximate the value of future actions through simulations. In particular, we combine additional sampling approximations with inference simplifications to make reasoning over the complex GBA-POMDP possible. Third, we discuss two sophisticated models that move away from the tabular representations that were assumed in previous work. We show how a Bayes-net representation allows us to model and learn structure in the dynamics of the system. This includes the usage of an intricate and targeted sampling scheme to avoid the collapse of the posterior approximation. Lastly, we show the practical use of the GBA-POMDP withBayes-Adaptive Deep Dropout Reinforcement learning (BADDr)'', which employs neural networks to model the dynamics of the environment. BADDr, with the expressiveness and scalability of neural networks, showcases how Bayesian reinforcement learning can be both principled and practical.--Author's abstract
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".