{"id":"W3119430014","doi":"10.65109/bwtd6573","title":"Partially Observable Mean Field Reinforcement Learning","year":2021,"lang":"en","type":"preprint","venue":"","topic":"Reinforcement Learning in Robotics","field":"Computer Science","cited_by":6,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Alberta; University of Waterloo","funders":"","keywords":"Reinforcement learning; Computer science; Scalability; Field (mathematics); Visibility; Artificial intelligence; Observable; Mathematical optimization; Set (abstract data type); Q-learning; Machine learning; Mathematics","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.002311522,0.001106707,0.001605287,0.000483692,0.000506138,0.0009643201,0.001999166,0.001516661,0.001902437],"category_scores_gemma":[0.00960112,0.0004747375,0.0005252878,0.0005041896,0.001836945,0.001417669,0.001233023,0.001941196,0.000329298],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.001494452,"about_ca_system_score_gemma":0.001496114,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.005155181,"about_ca_topic_score_gemma":0.004260678,"domain_scores_codex":[0.9988493,0.000478377,0.00004294546,0.0002424354,0.0002277218,0.0001593005],"domain_scores_gemma":[0.9931769,0.004833802,0.0006220611,0.0004447878,0.0005840245,0.000338357],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"simulation_or_modeling","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0000787707,0.00004592998,0.0005341673,0.00003688344,0.00002626808,0.00004413708,0.00003003138,0.9705999,0.0003689144,0.01443961,0.000690599,0.01310477],"study_design_scores_gemma":[0.00001136293,0.0000127716,0.00003579822,0.000002034748,0.000002138793,0.000003640445,0.000002023859,0.9936748,0.00009574612,0.00604828,0.0001089459,0.00000241604],"study_design_candidate":"simulation_or_modeling","study_design_consensus":"simulation_or_modeling","genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.03858088,0.0002390627,0.9572496,0.0004700775,0.00006993282,0.00006620095,0.00006457226,0.00057376,0.002686008],"genre_scores_gemma":[0.9287167,0.000112222,0.06788044,0.0001888983,0.00004745401,0.0001351826,0.00009553202,0.00004540374,0.002778096],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.005155181,"threshold_uncertainty_score":0.01222467,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.03734177779240893,"score_gpt":0.2667191034409659,"score_spread":0.229377325648557,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}