Who Should I Trust? Uncertainty and Risk for Knowledge Transfer from Multiple Sources in Reinforcement Learning Domains
Bibliographic record
Abstract
Despite the recent success of reinforcement learning (RL) in simulated domains and industrial applications, sample-efficiency remains a fundamental limitation of many model-free algorithms. Transfer learning mitigates this problem by using prior knowledge obtained by solving one set of tasks in order to accelerate the convergence on future tasks. However, while many frameworks have been proposed to successfully transfer different kinds of knowledge representations between tasks, existing transfer learning approaches remain largely incognizant to risk and uncertainty during transfer. In this thesis, we identify two sources of risk that must be addressed in order to make transfer learning from multiple knowledge sources more reliable and autonomous: epistemic uncertainty arises due to a lack of uncertainty about the model predictions, while aleatory uncertainty arises due to the stochastic nature of the environment. We address epistemic uncertainty by leveraging Bayesian model combination (BMC) to quantify and utilize uncertainty over the selection of knowledge sources for transfer, and we develop novel analytical techniques to efficiently tackle approximate Bayesian inference to train such models. We demonstrate the success of the proposed framework by transferring value functions, policies, and raw demonstrations between tasks. Next, we begin our treatment of aleatory uncertainty by highlighting some of the challenges in accounting for such risks during transfer, namely the lack of computationally tractable solutions that also provide theoretical assurances on the quality of transfer and control of risk. To mitigate this problem, we begin with two transfer learning approaches that have been highly successful in risk-neutral transfer -- namely potential-based reward shaping and successor features -- and extend them to the risk-sensitive setting. We empirically validate all our contributions on standard RL benchmarks, where they are shown to outperform other state-of-the-art transfer learning approaches in terms of robustness to noise and covariance shift in the training data, risk-sensitivity, and ease of interpretation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".