{"id":"W2197326744","doi":"10.1017/s0269888910000354","title":"Robotics competitions as benchmarks for AI research","year":2011,"lang":"en","type":"article","venue":"The Knowledge Engineering Review","topic":"Reinforcement Learning in Robotics","field":"Computer Science","cited_by":30,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Manitoba","funders":"","keywords":"Robotics; Benchmark (surveying); Artificial intelligence; Computer science; Applications of artificial intelligence; Robot","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":["metaresearch"],"consensus_categories":[],"category_scores_codex":[0.0289837,0.0005800533,0.0009266664,0.005247228,0.001675966,0.005217546,0.002015297,0.001575592,0.007618967],"category_scores_gemma":[0.03370391,0.0001972412,0.0003713329,0.004930849,0.002603533,0.003874397,0.003065137,0.001876201,0.001639322],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.005317928,"about_ca_system_score_gemma":0.003595746,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.004005329,"about_ca_topic_score_gemma":0.006203526,"domain_scores_codex":[0.9806823,0.01060141,0.0007933834,0.0008917141,0.006122302,0.0009090379],"domain_scores_gemma":[0.9614685,0.01602532,0.003519653,0.001340006,0.01464217,0.003004334],"domain_codex":null,"domain_gemma":"evaluation","domain_candidate":"evaluation","domain_consensus":null,"study_design_codex":"theoretical_or_conceptual","study_design_gemma":"theoretical_or_conceptual","study_design_scores_codex":[0.0007215535,0.0004717026,0.004384163,0.002092657,0.0001496134,0.0001412881,0.0008038656,0.005258766,0.0007261587,0.4566461,0.1747127,0.3538916],"study_design_scores_gemma":[0.0002494017,0.0008991752,0.01570738,0.002599057,0.0000748347,0.0001997229,0.002926347,0.004959241,0.001745924,0.1367934,0.8337182,0.0001272352],"study_design_candidate":"theoretical_or_conceptual","study_design_consensus":"theoretical_or_conceptual","genre_codex":"other","genre_gemma":"methods","genre_scores_codex":[0.0848432,0.3374326,0.01410193,0.09013498,0.01418706,0.0004680233,0.0012036,0.000293724,0.4573349],"genre_scores_gemma":[0.8602203,0.0717147,0.01385811,0.01525068,0.005757436,0.0006780049,0.002398675,0.000302354,0.0298197],"genre_candidate":"methods","genre_consensus":null,"teacher_disagreement_score":0.9710163,"threshold_uncertainty_score":0.1532823,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.1069341295820813,"score_gpt":0.3609366505106727,"score_spread":0.2540025209285914,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}