{"id":"W2593998405","doi":"10.1257/aer.20170299","title":"Evaluating Strategic Forecasters","year":2018,"lang":"en","type":"article","venue":"American Economic Review","topic":"Auction Theory and Applications","field":"Decision Sciences","cited_by":18,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"","keywords":"Generality; Robustness (evolution); Economics; Principal (computer security); Quality (philosophy); Mechanism design; Microeconomics; Econometrics; Simple (philosophy); Computer science; Mathematical economics; Computer security","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":["insufficient_payload"],"consensus_categories":["insufficient_payload"],"category_scores_codex":[0.002737006,0.00008984036,0.000316465,0.00004985432,0.000140317,0.0000677737,0.0004891334,0.000009951157,0.005507422],"category_scores_gemma":[0.0002644207,0.00006769244,0.0001088087,0.0003081235,0.0004938738,0.0001554441,0.00005121552,0.00005026129,0.01328174],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0000416968,"about_ca_system_score_gemma":0.00007670647,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.00001835825,"about_ca_topic_score_gemma":0.00001536389,"domain_scores_codex":[0.9985437,0.0001926442,0.0005970399,0.0003681177,0.0001483475,0.0001501682],"domain_scores_gemma":[0.9982697,0.00037004,0.0005608322,0.000640893,0.00007645485,0.00008204373],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"not_applicable","study_design_scores_codex":[0.000002958726,0.000005934659,0.0001395089,0.000006611648,0.000008121347,1.000331e-7,0.00003465264,0.00001328668,0.00002460132,0.01869428,0.009243686,0.9718263],"study_design_scores_gemma":[0.00019958,0.0004901824,0.0008291941,0.0002670789,0.00006197861,0.00004931853,0.001539613,0.005311595,0.0001368804,0.205101,0.7855882,0.0004253439],"study_design_candidate":"design_other","study_design_consensus":null,"genre_codex":"empirical","genre_gemma":"empirical","genre_scores_codex":[0.5493724,0.007589113,0.006152159,0.01057396,0.0008140236,0.001171616,0.00002638827,0.0001286913,0.4241717],"genre_scores_gemma":[0.9860969,0.002790996,0.002558545,0.005352112,0.0003508904,0.0000789038,0.000003011589,0.00001196426,0.00275663],"genre_candidate":"empirical","genre_consensus":"empirical","teacher_disagreement_score":0.9714009,"threshold_uncertainty_score":0.9954017,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.4133607625397049,"score_gpt":0.5322415473275208,"score_spread":0.1188807847878159,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}