{"id":"W4401340787","doi":"10.1017/s0003055424000716","title":"Improving Probabilistic Models In Text Classification Via Active Learning","year":2024,"lang":"en","type":"article","venue":"American Political Science Review","topic":"Topic Modeling","field":"Computer Science","cited_by":2,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"","keywords":"Probabilistic logic; Computer science; Artificial intelligence; Machine learning; Active learning (machine learning)","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.001332423,0.0001429496,0.0002814403,0.0001806216,0.0001119565,0.0002216234,0.0009554132,0.00002041566,0.000008721549],"category_scores_gemma":[0.0007715548,0.0001170112,0.00006419267,0.002571185,0.0006818705,0.001202909,0.000279266,0.0003454807,0.00007653742],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0006389385,"about_ca_system_score_gemma":0.0006234803,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.0008165838,"about_ca_topic_score_gemma":0.00000863293,"domain_scores_codex":[0.9972344,0.0001353756,0.0003746945,0.0008296436,0.0005096323,0.0009162367],"domain_scores_gemma":[0.9988031,0.0002389668,0.00007356575,0.0004750175,0.00008650951,0.0003228787],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[1.523391e-7,0.000008415072,0.000009602577,0.000185101,6.523752e-7,0.000003440766,0.00005465612,0.0000860587,0.0001737005,0.4824961,0.000001939705,0.5169802],"study_design_scores_gemma":[0.00002132992,0.00004094795,0.0004060247,0.0009386343,0.000007205252,0.00002008442,0.00004330761,0.9676234,0.00003409899,0.03010357,0.0005986861,0.0001627479],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.003001628,0.004141355,0.9755214,0.009504023,0.0001498992,0.0004715704,6.93547e-7,0.0002753483,0.006934019],"genre_scores_gemma":[0.9861433,0.0006805978,0.01210442,0.0009293805,0.0000397637,0.00006104643,4.674856e-7,0.000007895767,0.00003314202],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.9831417,"threshold_uncertainty_score":0.4771577,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.03848909323123417,"score_gpt":0.328479701192705,"score_spread":0.2899906079614708,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}