{"id":"W4407031081","doi":"10.32388/8g8tb2","title":"Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey","year":2025,"lang":"en","type":"preprint","venue":"Qeios","topic":"Software Engineering Research","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"","keywords":"Reinforcement learning; Computer science; Compiler; Code generation; Code (set theory); Resource allocation; Resource (disambiguation); Artificial intelligence; Programming language; Computer security","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.002572268,0.0009917638,0.0007890468,0.001535663,0.0003670359,0.001279756,0.001786554,0.001052344,0.002544339],"category_scores_gemma":[0.01037458,0.0006974357,0.0008934839,0.001676843,0.001175148,0.001784458,0.001216098,0.001564736,0.001437617],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0008853382,"about_ca_system_score_gemma":0.001645208,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.002003061,"about_ca_topic_score_gemma":0.001623724,"domain_scores_codex":[0.998243,0.0005449657,0.0001480169,0.0003100582,0.000664908,0.00008916842],"domain_scores_gemma":[0.9938369,0.004343194,0.0002531305,0.0007863471,0.0006775526,0.0001027522],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"not_applicable","study_design_scores_codex":[0.00005509874,0.0001401338,0.002037632,0.001966067,0.00005078069,0.00005875728,0.0002267345,0.07655607,0.005563887,0.02431264,0.003949814,0.8850824],"study_design_scores_gemma":[0.00008675765,0.0006374312,0.002358335,0.001659296,0.0001207305,0.0007633731,0.0002890629,0.6807299,0.03223961,0.06851258,0.2124731,0.000129914],"study_design_candidate":"not_applicable","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"review","genre_scores_codex":[0.02170406,0.09660996,0.8596317,0.002023923,0.0001816251,0.0002242341,0.0001864011,0.003899074,0.01553908],"genre_scores_gemma":[0.2149146,0.1004181,0.6734844,0.0009304134,0.0003758024,0.0004093819,0.0008108813,0.001904655,0.006751769],"genre_candidate":"review","genre_consensus":null,"teacher_disagreement_score":0.002572268,"threshold_uncertainty_score":0.01360363,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.04221355533285516,"score_gpt":0.2985258908404057,"score_spread":0.2563123355075506,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}