MétaCan
Menu
Back to cohort

Toward Off-Policy Learning Control with Function Approximation

2010· article· en· W13294968 on OpenAlexaff
Hamid Reza Maei, Csaba Szepesv ri, Shalabh Bhatnagar, Richard Sutton

Bibliographic record

VenueInternational Conference on Machine Learning · 2010
Typearticle
Languageen
FieldComputer Science
TopicReinforcement Learning in Robotics
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsTemporal difference learningFunction approximationApproximation algorithmFunction (biology)Computer scienceLinear approximationBellman equationOptimal controlExtension (predicate logic)Mathematical optimizationControl (management)Artificial intelligenceMathematicsReinforcement learningArtificial neural networkNonlinear system

Abstract

fetched live from OpenAlex

We present the first temporal-difference learning algorithm for off-policy control with unrestricted linear function approximation whose per-time-step complexity is linear in the number of features. Our algorithm, Greedy-GQ, is an extension of recent work on gradient temporal-difference learning, which has hitherto been restricted to a prediction (policy evaluation) setting, to a control setting in which the target policy is greedy with respect to a linear approximation to the optimal action-value function. A limitation of our control setting is that we require the behavior policy to be stationary. We call this setting latent learning because the optimal policy, though learned, is not manifest in behavior. Popular off-policy algorithms such as Q-learning are known to be unstable in this setting when used with linear function approximation. In reinforcement learning, the term “off-policy learning” refers to learning about one way of behaving, called the target policy, from data generated by another way of selecting actions, called the behavior policy. The target policy is often an approximation to the optimal policy, which is typically deterministic, whereas the behavior policy is often stochastic, exploring all possible actions in each state as part of finding the optimal policy. Freeing the behavior policy from the target policy enables a greater variety of exploration strategies to be used. It also enables learning from training data generated by unrelated controllers, including manual human control, and from previously collected data. A third reason for interest in off-policy learning is that it permits learning about multiple target policies (e.g., optimal policies for multiple subgoals) from a single stream of data generated by a

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.006
Threshold uncertainty score0.016

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0030.008
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0010.001
Science and technology studies0.0010.002
Scholarly communication0.0020.002
Open science0.0020.003
Research integrity0.0020.004
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.023
GPT teacher head0.271
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations175
Published2010
Admission routes1
Has abstractyes

Explore more

Same venueInternational Conference on Machine LearningSame topicReinforcement Learning in RoboticsFrench-language works237,207