{"id":"W3009319801","doi":"10.1145/3328778.3372601","title":"Analyzing CS1 Student Code Using Code Embeddings","year":2020,"lang":"en","type":"article","venue":"","topic":"Software Engineering Research","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"","keywords":"Computer science; Code (set theory); Programming language; Source code; Theoretical computer science; Parallel computing","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.001488682,0.0008173301,0.0003552194,0.002186015,0.0003977481,0.001251498,0.0006896142,0.0007996925,0.00179653],"category_scores_gemma":[0.01864592,0.0002146318,0.000514138,0.001930763,0.000622229,0.001969124,0.001238612,0.001391239,0.0007520976],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00079493,"about_ca_system_score_gemma":0.0008648435,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.001969281,"about_ca_topic_score_gemma":0.0031299,"domain_scores_codex":[0.9982091,0.0005477099,0.0001229193,0.0004034473,0.0005797443,0.0001370804],"domain_scores_gemma":[0.9888842,0.005063366,0.001537362,0.001167502,0.002873229,0.0004743629],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"bench_or_experimental","study_design_scores_codex":[0.0009664812,0.001032682,0.1407159,0.0005584228,0.0001694614,0.000641634,0.002185261,0.191927,0.02211341,0.01817476,0.0181792,0.6033358],"study_design_scores_gemma":[0.00002750893,0.0002840974,0.01341693,0.00004936185,0.00001939828,0.0001958974,0.000570571,0.9429923,0.01218293,0.02528985,0.004920921,0.00005014124],"study_design_candidate":"bench_or_experimental","study_design_consensus":null,"genre_codex":"empirical","genre_gemma":"empirical","genre_scores_codex":[0.6737003,0.000365984,0.3165021,0.0008431631,0.0001463991,0.0001264313,0.002296861,0.003342811,0.002675978],"genre_scores_gemma":[0.9102204,0.0001212941,0.08272339,0.00007512709,0.0000424532,0.0001305262,0.003664102,0.0002985499,0.00272408],"genre_candidate":"empirical","genre_consensus":"empirical","teacher_disagreement_score":0.002186015,"threshold_uncertainty_score":0.007872999,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.05509329753359742,"score_gpt":0.3399035939572613,"score_spread":0.2848102964236639,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}