Who Should I Trust? Uncertainty and Risk for Knowledge Transfer from Multiple Sources in Reinforcement Learning Domains
Notice bibliographique
Résumé
Despite the recent success of reinforcement learning (RL) in simulated domains and industrial applications, sample-efficiency remains a fundamental limitation of many model-free algorithms. Transfer learning mitigates this problem by using prior knowledge obtained by solving one set of tasks in order to accelerate the convergence on future tasks. However, while many frameworks have been proposed to successfully transfer different kinds of knowledge representations between tasks, existing transfer learning approaches remain largely incognizant to risk and uncertainty during transfer. In this thesis, we identify two sources of risk that must be addressed in order to make transfer learning from multiple knowledge sources more reliable and autonomous: epistemic uncertainty arises due to a lack of uncertainty about the model predictions, while aleatory uncertainty arises due to the stochastic nature of the environment. We address epistemic uncertainty by leveraging Bayesian model combination (BMC) to quantify and utilize uncertainty over the selection of knowledge sources for transfer, and we develop novel analytical techniques to efficiently tackle approximate Bayesian inference to train such models. We demonstrate the success of the proposed framework by transferring value functions, policies, and raw demonstrations between tasks. Next, we begin our treatment of aleatory uncertainty by highlighting some of the challenges in accounting for such risks during transfer, namely the lack of computationally tractable solutions that also provide theoretical assurances on the quality of transfer and control of risk. To mitigate this problem, we begin with two transfer learning approaches that have been highly successful in risk-neutral transfer -- namely potential-based reward shaping and successor features -- and extend them to the risk-sensitive setting. We empirically validate all our contributions on standard RL benchmarks, where they are shown to outperform other state-of-the-art transfer learning approaches in terms of robustness to noise and covariance shift in the training data, risk-sensitivity, and ease of interpretation.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».