{"id":"W4393213239","doi":"10.1145/3597503.3639194","title":"ChatGPT Incorrectness Detection in Software Reviews","year":2024,"lang":"en","type":"preprint","venue":"","topic":"Software Engineering Research","field":"Computer Science","cited_by":10,"is_retracted":false,"has_abstract":true,"ca_institutions":"York University; University of Calgary","funders":"Natural Sciences and Engineering Research Council of Canada","keywords":"Computer science; Suite; Benchmark (surveying); Selection (genetic algorithm); Generative grammar; Artificial intelligence; Software; Machine learning; Natural language processing; Information retrieval; Programming language","routes":{"ca_aff":true,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":["metaresearch"],"consensus_categories":[],"category_scores_codex":[0.01700345,0.0009707248,0.001175379,0.008435271,0.0009967615,0.001569885,0.00138513,0.00138896,0.0009627976],"category_scores_gemma":[0.2054146,0.0005225794,0.0004825815,0.004014659,0.0005735773,0.002514509,0.001883352,0.001143089,0.000860077],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0009971063,"about_ca_system_score_gemma":0.0009908275,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.002053674,"about_ca_topic_score_gemma":0.003936095,"domain_scores_codex":[0.9509097,0.02208221,0.004356717,0.00537654,0.01611172,0.001163087],"domain_scores_gemma":[0.6284948,0.2750643,0.04048543,0.01073531,0.04291495,0.002305258],"domain_codex":null,"domain_gemma":"evaluation","domain_candidate":"evaluation","domain_consensus":null,"study_design_codex":"observational","study_design_gemma":"observational","study_design_scores_codex":[0.001050873,0.0005272672,0.54755,0.003663348,0.0003462513,0.002194252,0.02057349,0.004798807,0.02845288,0.001399903,0.01305697,0.376386],"study_design_scores_gemma":[0.0001474544,0.001756516,0.7336906,0.001362072,0.0004589212,0.006116858,0.01020063,0.1482521,0.05659498,0.003904062,0.03711164,0.0004041562],"study_design_candidate":"observational","study_design_consensus":"observational","genre_codex":"empirical","genre_gemma":"empirical","genre_scores_codex":[0.9606348,0.001753629,0.0299945,0.0004583587,0.0001287622,0.0004626432,0.001198502,0.002503237,0.00286561],"genre_scores_gemma":[0.9705085,0.0003996445,0.02489638,0.0002288742,0.00006553668,0.0002819729,0.001953646,0.0001969128,0.001468543],"genre_candidate":"empirical","genre_consensus":"empirical","teacher_disagreement_score":0.9829966,"threshold_uncertainty_score":0.08992392,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.02675092483946025,"score_gpt":0.299219897492254,"score_spread":0.2724689726527938,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}