{"id":"W4399567378","doi":"10.1145/3650105.3652299","title":"Fine Tuning Large Language Model for Secure Code Generation","year":2024,"lang":"en","type":"article","venue":"","topic":"Software Engineering Research","field":"Computer Science","cited_by":23,"is_retracted":false,"has_abstract":true,"ca_institutions":"Queen's University; Concordia University","funders":"","keywords":"Computer science; Code (set theory); Code generation; Vulnerability (computing); Face (sociological concept); Source code; Programming language; Computer security","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0002519592,0.00005737807,0.00004829355,0.00008010309,0.00004235818,0.0001956904,0.0002532752,0.00003173478,0.00001683048],"category_scores_gemma":[0.0001218362,0.0000497763,0.00002999373,0.0001838644,0.000003207262,0.0002296678,0.00009554324,0.00007997575,0.00002979472],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00002969167,"about_ca_system_score_gemma":0.00005046981,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.000003353792,"about_ca_topic_score_gemma":0.00003771319,"domain_scores_codex":[0.9993784,0.000006216927,0.00006741298,0.0002093085,0.0001428944,0.000195717],"domain_scores_gemma":[0.9995466,0.0001575615,0.000003963217,0.0002184385,0.00003286517,0.00004054196],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"theoretical_or_conceptual","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.000003273926,0.00005460013,0.0001702066,0.0003083629,0.00005539635,0.00006886869,0.0102482,0.167858,0.06243912,0.5738832,0.1380498,0.04686091],"study_design_scores_gemma":[0.00007212556,0.00001224825,0.00001207275,0.000009328568,0.000001075862,0.000002855292,0.000004946971,0.9945921,0.002886614,0.0002613462,0.002080578,0.00006470773],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.01182767,0.0003806213,0.986114,0.0005731352,0.0001972142,0.0001082398,0.00001270537,0.0006808678,0.0001055113],"genre_scores_gemma":[0.8112856,0.000001666339,0.1830961,0.00006573687,0.0001514186,0.00003632683,0.00001538552,0.00001245482,0.005335362],"genre_candidate":"methods","genre_consensus":null,"teacher_disagreement_score":0.8267341,"threshold_uncertainty_score":0.2029818,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.03353709414965798,"score_gpt":0.312070987112272,"score_spread":0.278533892962614,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}