Towards better generalization capabilities of reinforcement learning agents via self-supervision
Bibliographic record
Abstract
The notion of general intelligence of an agent is closely linked with its ability to solve different tasks, given a certain amount of prior information.However, the recent success of deep reinforcement learning agents in challenging domains such as video games and robotic control has been hindered by the limited ability to re-use pre-trained agents on unseen tasks.Since reinforcement learning tasks are primarily specified through the reward signal, incorporating additional reward-free learning objectives that better capture other components of the environment should lead to task-agnostic representations which generalize better.This thesis studies the idea of using self-supervised and unsupervised learning to discover state representations for better generalization of reinforcement learning agents.First, I propose a representation learning framework for projecting policies onto a Reproducing Kernel Hilbert Space, which comes with theoretical guarantees and variance-reduction properties.In order to alleviate the exponential space complexity of the RKHS framework, I then develop a method to capture multi-step behavioral similarities from trajectories, which helps improve zero-shot generalization performance on hard procedural generalization tasks.Third, I show how the auxiliary task of predicting future states via the InfoMax principle [Linsker, 1988] can help the agent learn under a distribution shift.Finally, I present a first attempt to learn generalizable state representations from offline high-dimensional datasets using contrastive learning across multiple generalized value functions.Thank you to my partner Kayla, my parents and my sister for their support and encouragement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".