{"id":"W4403935894","doi":"10.1145/3652620.3676878","title":"Towards Model Repair by Human Opinion--Guided Reinforcement Learning","year":2024,"lang":"en","type":"article","venue":"","topic":"Reinforcement Learning in Robotics","field":"Computer Science","cited_by":2,"is_retracted":false,"has_abstract":true,"ca_institutions":"McMaster University","funders":"","keywords":"Reinforcement learning; Computer science; Reinforcement; Artificial intelligence; Human–computer interaction; Engineering; Structural engineering","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.002038277,0.001157165,0.001013107,0.0005555305,0.0003830891,0.0009037158,0.001500494,0.001365968,0.001871002],"category_scores_gemma":[0.01081493,0.0004333424,0.0006124919,0.0003114871,0.001075211,0.001274138,0.001595292,0.001723035,0.0004110621],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0008110037,"about_ca_system_score_gemma":0.001197819,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.004364211,"about_ca_topic_score_gemma":0.004416667,"domain_scores_codex":[0.9990886,0.000386468,0.00003779851,0.00023083,0.0001590499,0.0000972557],"domain_scores_gemma":[0.9946679,0.003611973,0.0006076425,0.0004022751,0.0004604449,0.0002497377],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"simulation_or_modeling","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0001511984,0.000139492,0.002686755,0.0001265147,0.00006844746,0.000132234,0.0003535889,0.894218,0.00386108,0.01068361,0.00184937,0.08572979],"study_design_scores_gemma":[0.00001037147,0.00002172701,0.00006414585,0.000004906472,0.000004950291,0.000008896572,0.00001235862,0.994783,0.0003409302,0.004544921,0.0001999982,0.000003776945],"study_design_candidate":"simulation_or_modeling","study_design_consensus":"simulation_or_modeling","genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.04137831,0.0001667786,0.9556397,0.0003393627,0.00003344787,0.00004676125,0.00003304076,0.0005909911,0.001771586],"genre_scores_gemma":[0.8697605,0.00009122123,0.128422,0.0002297027,0.00004641921,0.0001001369,0.0001025663,0.0001003887,0.001146999],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.004364211,"threshold_uncertainty_score":0.01077956,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.03763230770512827,"score_gpt":0.3118472252698286,"score_spread":0.2742149175647003,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}