{"id":"W2963241005","doi":"10.1145/3322640.3326711","title":"A Reliable and Accurate Multiple Choice Question Answering System for Due Diligence","year":2019,"lang":"en","type":"article","venue":"","topic":"Topic Modeling","field":"Computer Science","cited_by":4,"is_retracted":false,"has_abstract":true,"ca_institutions":"Cisco Systems (Canada)","funders":"","keywords":"Question answering; Computer science; Due diligence; Scarcity; Artificial intelligence; Classifier (UML); Task (project management); Machine learning; Oversampling; Bandwidth (computing); Finance","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0002311592,0.00007432653,0.0001024107,0.00003258921,0.00006254805,0.0001160601,0.000257604,0.00003691308,0.000001606352],"category_scores_gemma":[0.00005298889,0.00006570151,0.00001848243,0.00007942106,0.000004987894,0.0005200816,0.0001185775,0.00004090749,0.00001869721],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00003240356,"about_ca_system_score_gemma":0.00001898869,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.0003169595,"about_ca_topic_score_gemma":0.00003565858,"domain_scores_codex":[0.999262,0.00001019971,0.0001410014,0.000323229,0.00008812033,0.0001754606],"domain_scores_gemma":[0.9993706,0.0001440139,0.00004276822,0.000343415,0.00005237816,0.00004677966],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"theoretical_or_conceptual","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.00004462807,0.00008112536,0.2009083,0.002538414,0.00005192503,0.0000130742,0.002215129,0.08409472,0.04926089,0.5986215,0.0003706748,0.06179966],"study_design_scores_gemma":[0.0002140412,0.00002542923,0.003008675,0.0000666128,0.000001695668,0.000008296129,0.00003401305,0.9917454,0.003713302,0.0001327548,0.0009524172,0.00009740503],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.2885728,0.00005124859,0.7100106,0.000114642,0.0003363875,0.0002548468,3.758776e-7,0.0001739775,0.0004850784],"genre_scores_gemma":[0.9137626,0.000002510781,0.08543131,0.00004628214,0.00003746856,0.00002341449,4.496002e-7,0.00000480855,0.0006911143],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.9076506,"threshold_uncertainty_score":0.267923,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.02117962379379888,"score_gpt":0.2575918210204401,"score_spread":0.2364121972266412,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}