{"id":"W2333065964","doi":"10.1145/2854946.2854993","title":"Are Secondary Assessors Uncertain When They Disagree About Relevance Judgements?","year":2016,"lang":"en","type":"article","venue":"","topic":"Mobile Crowdsensing and Crowdsourcing","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Waterloo","funders":"Natural Sciences and Engineering Research Council of Canada; King Saud bin Abdulaziz University for Health Science; University of Waterloo; King Abdulaziz University; King Saud University","keywords":"Judgement; Relevance (law); Certainty; Adjudication; Test (biology); Information retrieval; Psychology; Computer science; Mathematics; Political science; Law","routes":{"ca_aff":true,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0004237174,0.0002155929,0.0002212445,0.00007692558,0.0002050589,0.0002576611,0.000926792,0.00007594034,0.0002018202],"category_scores_gemma":[0.0001940606,0.0001345617,0.0000931285,0.0001348095,0.00006924028,0.0007235561,0.0003363827,0.0001353019,0.000362495],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00008703404,"about_ca_system_score_gemma":0.00006592318,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.00007286653,"about_ca_topic_score_gemma":0.0001807115,"domain_scores_codex":[0.9981493,0.0001010981,0.0002878491,0.0006198413,0.0003525933,0.000489284],"domain_scores_gemma":[0.9981442,0.0002934663,0.000214607,0.001095416,0.0001066936,0.0001456196],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"not_applicable","study_design_scores_codex":[0.00001802538,0.0001814918,0.01967449,0.00006652156,0.00009245155,0.0001434206,0.001483021,0.00008989355,0.007953366,0.04582916,0.08111097,0.8433572],"study_design_scores_gemma":[0.006961695,0.0003998718,0.1638024,0.002304667,0.0000670337,0.0001971501,0.001205939,0.01483491,0.03493508,0.1688247,0.6025754,0.003891201],"study_design_candidate":"design_other","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.1782133,0.0002739237,0.7585139,0.008493279,0.001211139,0.000287246,0.000009858379,0.0009265716,0.05207081],"genre_scores_gemma":[0.9642867,0.0000287705,0.01657305,0.001782653,0.0001348855,0.0000157021,7.865947e-7,0.00002132707,0.01715609],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.839466,"threshold_uncertainty_score":0.5487266,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.02126795997712938,"score_gpt":0.2489718051851484,"score_spread":0.227703845208019,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}