{"id":"W2517079402","doi":"10.1145/2960310.2960337","title":"Benchmarking Introductory Programming Exams","year":2016,"lang":"en","type":"article","venue":"","topic":"Teaching and Learning Programming","field":"Computer Science","cited_by":8,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"","keywords":"Benchmarking; Computer science; Set (abstract data type); Mathematics education; Programming language; Psychology; Management","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.01567963,0.000948254,0.0009472214,0.008089972,0.0008157783,0.003094879,0.001598026,0.001088128,0.005861958],"category_scores_gemma":[0.1127926,0.0003146827,0.000789554,0.007815493,0.0006288601,0.001870992,0.002809059,0.001301771,0.003346231],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00161148,"about_ca_system_score_gemma":0.001181249,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.002790609,"about_ca_topic_score_gemma":0.002997063,"domain_scores_codex":[0.9641674,0.01208478,0.004610931,0.004394892,0.0118493,0.002892681],"domain_scores_gemma":[0.8666236,0.03822078,0.01464371,0.01560154,0.05735831,0.007552004],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"observational","study_design_gemma":"observational","study_design_scores_codex":[0.001440914,0.001754033,0.5807994,0.0005217198,0.0004395563,0.0002613129,0.0024227,0.01342385,0.008016876,0.00650146,0.01672261,0.3676956],"study_design_scores_gemma":[0.00006030135,0.001490649,0.9478696,0.0001076944,0.00008402199,0.0002329977,0.0012284,0.008955699,0.01421655,0.002238744,0.02343114,0.00008420658],"study_design_candidate":"observational","study_design_consensus":"observational","genre_codex":"empirical","genre_gemma":"empirical","genre_scores_codex":[0.9559168,0.0007325064,0.01267722,0.0002884664,0.0002568099,0.0004019011,0.003730883,0.0007661345,0.02522911],"genre_scores_gemma":[0.977711,0.0002108268,0.008754684,0.00008677995,0.0000926669,0.0002211186,0.008602998,0.000199713,0.004120234],"genre_candidate":"empirical","genre_consensus":"empirical","teacher_disagreement_score":0.01567963,"threshold_uncertainty_score":0.08292282,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.01227547792997245,"score_gpt":0.2325014494734201,"score_spread":0.2202259715434476,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}