{"id":"W4388761312","doi":"10.17760/d20486919","title":"Bayesian partially observable reinforcement learning","year":2023,"lang":"en","type":"dissertation","venue":"","topic":"Reinforcement Learning in Robotics","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"Science North","funders":"","keywords":"Computer science; Reinforcement learning; Inference; State (computer science); Bayesian inference; Artificial intelligence; Bayesian probability","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.002115515,0.001409869,0.002265867,0.0007234439,0.0005958902,0.001462874,0.00214481,0.002112748,0.005781085],"category_scores_gemma":[0.01069745,0.0007594335,0.0006404057,0.0007346419,0.001660424,0.001903733,0.001529336,0.002342962,0.001014071],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.002048407,"about_ca_system_score_gemma":0.002154075,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.01064532,"about_ca_topic_score_gemma":0.01123096,"domain_scores_codex":[0.998489,0.0006818589,0.00006610605,0.0002989877,0.0002695812,0.0001943899],"domain_scores_gemma":[0.9939931,0.004361428,0.000437641,0.0003034575,0.000587791,0.0003165438],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"simulation_or_modeling","study_design_gemma":"theoretical_or_conceptual","study_design_scores_codex":[0.0002069142,0.0001048183,0.001628048,0.0001584724,0.00007823301,0.0001416827,0.0001066428,0.8927065,0.0003054691,0.0593718,0.004042673,0.04114872],"study_design_scores_gemma":[0.00003491014,0.00002089749,0.0001382384,0.00001314037,0.000008769108,0.00001227131,0.000008775834,0.9643597,0.00008099969,0.03448397,0.0008287118,0.00000960482],"study_design_candidate":"theoretical_or_conceptual","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.02688671,0.001111968,0.9551683,0.001351619,0.0001786019,0.0001481839,0.0005195152,0.001139443,0.0134956],"genre_scores_gemma":[0.8833604,0.0006230904,0.1039815,0.0005080828,0.0001272012,0.0003585038,0.0007371681,0.0001017252,0.01020222],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.01064532,"threshold_uncertainty_score":0.02116674,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.02529112251535362,"score_gpt":0.2726790296713985,"score_spread":0.2473879071560448,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}