{"id":"W2050712820","doi":"10.1145/775047.775138","title":"Discovering word senses from text","year":2002,"lang":"en","type":"article","venue":"","topic":"Natural Language Processing Techniques","field":"Computer Science","cited_by":597,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Alberta","funders":"Natural Sciences and Engineering Research Council of Canada","keywords":"Word (group theory); Computer science; Cluster analysis; Precision and recall; Set (abstract data type); Similarity (geometry); Centroid; Artificial intelligence; Natural language processing; Feature (linguistics); Feature vector; Element (criminal law); Space (punctuation); Cluster (spacecraft); Domain (mathematical analysis); Recall; Information retrieval; Mathematics; Linguistics; Image (mathematics)","routes":{"ca_aff":true,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.001752223,0.001219283,0.001170246,0.01320937,0.001387881,0.002262923,0.001090056,0.001110854,0.001652997],"category_scores_gemma":[0.01469443,0.0005952127,0.001394505,0.005855973,0.0009647218,0.004321162,0.002165837,0.0008536202,0.001658181],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0006646673,"about_ca_system_score_gemma":0.001421218,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.001935046,"about_ca_topic_score_gemma":0.004160596,"domain_scores_codex":[0.9962621,0.0008185006,0.0005463578,0.001278603,0.0009269146,0.0001674495],"domain_scores_gemma":[0.9916223,0.004516665,0.0009338977,0.001051967,0.001658318,0.0002167939],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"bench_or_experimental","study_design_scores_codex":[0.0007029556,0.0003875187,0.05773765,0.00245043,0.000482518,0.002354156,0.004287329,0.01106944,0.06191456,0.014102,0.02143087,0.8230805],"study_design_scores_gemma":[0.0002551924,0.0008032661,0.08306502,0.001238367,0.0009964285,0.01125989,0.01255222,0.433298,0.1533567,0.1327879,0.1698433,0.0005437078],"study_design_candidate":"bench_or_experimental","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.420488,0.003555354,0.5399799,0.000856318,0.0003531371,0.001040307,0.01650346,0.007489074,0.009734431],"genre_scores_gemma":[0.356296,0.001437975,0.6149664,0.00021273,0.0001546019,0.0005830666,0.0235802,0.0005921966,0.002176776],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.01320937,"threshold_uncertainty_score":0.009266794,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.01633319923786609,"score_gpt":0.2386434594540469,"score_spread":0.2223102602161808,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}