{"id":"W3119636502","doi":"10.18653/v1/2021.acl-long.569","title":"UnNatural Language Inference","year":2021,"lang":"en","type":"preprint","venue":"","topic":"Topic Modeling","field":"Computer Science","cited_by":5,"is_retracted":false,"has_abstract":true,"ca_institutions":"Mila - Quebec Artificial Intelligence Institute; McGill University","funders":"McGill University","keywords":"Computer science; Inference; Natural language processing; Artificial intelligence; Mandarin Chinese; Syntax; Suite; Transformer; Language model; Word order; Linguistics; History","routes":{"ca_aff":true,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0001031996,0.0001406066,0.0001669185,0.00005127659,0.0000255643,0.0004222534,0.001341179,0.0001556715,0.00009788724],"category_scores_gemma":[0.00005769261,0.00012539,0.00007856522,0.00008718317,0.00001017013,0.0001561653,0.003571471,0.0005151778,0.00003013601],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00003362595,"about_ca_system_score_gemma":0.0001995147,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.0003062277,"about_ca_topic_score_gemma":0.00006722499,"domain_scores_codex":[0.9988692,0.00003788226,0.000162054,0.0005311234,0.0002177589,0.0001819572],"domain_scores_gemma":[0.9985658,0.00004254012,0.00005353252,0.001210903,0.00006974252,0.00005746549],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.000001743275,0.000121225,0.001462658,0.0004107814,0.0001292713,0.0007393843,0.02154392,0.01705909,0.002089329,0.4164288,0.001396556,0.5386173],"study_design_scores_gemma":[0.00012769,0.000007804582,0.001105557,0.0001573702,0.000006652632,0.00001468435,0.0002446067,0.9873161,0.003025017,0.006991353,0.0004435016,0.0005596131],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.05886471,0.0005933785,0.9261444,0.0007892667,0.001185243,0.00007147397,6.934255e-7,0.0003126067,0.01203821],"genre_scores_gemma":[0.7546568,0.00001353899,0.243336,0.0004405303,0.00008832862,0.000006615456,0.000007291018,0.000004212489,0.001446719],"genre_candidate":"methods","genre_consensus":null,"teacher_disagreement_score":0.970257,"threshold_uncertainty_score":0.5113258,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.02933943172557949,"score_gpt":0.2968036872332471,"score_spread":0.2674642555076676,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}