{"id":"W4399158753","doi":"10.1075/kl.00008.par","title":"Word segmentation granularity in Korean","year":2024,"lang":"en","type":"article","venue":"Korean Linguistics","topic":"Natural Language Processing Techniques","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of British Columbia","funders":"","keywords":"Granularity; Segmentation; Natural language processing; Word (group theory); Computer science; Text segmentation; Linguistics; Artificial intelligence; Philosophy","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0003346729,0.0001309434,0.0001178436,0.0002200537,0.000049329,0.0003427749,0.0005233762,0.00007825901,0.000005285452],"category_scores_gemma":[0.0007044919,0.0001205202,0.00003915165,0.0006048705,0.00003747538,0.0001114433,0.0001520812,0.00029913,0.0000244413],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0001049962,"about_ca_system_score_gemma":0.00007317284,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.0001098357,"about_ca_topic_score_gemma":0.00005290143,"domain_scores_codex":[0.9989452,0.00003771643,0.000223441,0.0003451337,0.00022538,0.0002231238],"domain_scores_gemma":[0.9993823,0.00009967361,0.00003474256,0.0003246616,0.0001046661,0.00005394257],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"theoretical_or_conceptual","study_design_gemma":"theoretical_or_conceptual","study_design_scores_codex":[0.000004484951,0.00005464047,0.0009539324,0.0001676521,0.00001079134,0.000579964,0.001755434,0.000009941464,0.0007974898,0.8597411,0.00329642,0.1326282],"study_design_scores_gemma":[0.0002286627,0.00006807099,0.0003338142,0.000483257,0.00001972796,0.00002968693,0.00003600423,0.08430368,0.03979891,0.860837,0.01331905,0.0005420867],"study_design_candidate":"theoretical_or_conceptual","study_design_consensus":"theoretical_or_conceptual","genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.002965508,0.005362223,0.9769546,0.0005989872,0.002591401,0.0003018157,0.00001348571,0.002724794,0.008487155],"genre_scores_gemma":[0.5670502,0.00001755732,0.4323722,0.0001429779,0.0002844542,0.000006206279,0.00001408621,0.00001163597,0.0001007025],"genre_candidate":"methods","genre_consensus":null,"teacher_disagreement_score":0.5640847,"threshold_uncertainty_score":0.4914672,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.01480372139418494,"score_gpt":0.2943217223496563,"score_spread":0.2795180009554714,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}