{"id":"W2133503566","doi":"10.3115/1642011.1642021","title":"An unsupervised model for text message normalization","year":2009,"lang":"en","type":"article","venue":"","topic":"Digital Communication and Language","field":"Computer Science","cited_by":132,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"Natural Sciences and Engineering Research Council of Canada; University of Toronto","keywords":"Normalization (sociology); Computer science; Natural language processing; Artificial intelligence; Text messaging; Text message; Phone; Set (abstract data type); Test set; Construct (python library); Word (group theory); Variety (cybernetics); World Wide Web; Linguistics","routes":{"ca_aff":true,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.001999187,0.001256756,0.001259292,0.001336104,0.0006855449,0.001318255,0.002624133,0.001439895,0.003109433],"category_scores_gemma":[0.008283008,0.0006024204,0.001372688,0.001406198,0.001244932,0.002990532,0.001296167,0.002460976,0.002810932],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.001343711,"about_ca_system_score_gemma":0.001527984,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.004715025,"about_ca_topic_score_gemma":0.005768311,"domain_scores_codex":[0.9981306,0.0006227492,0.0001197926,0.0006443886,0.0003362311,0.000146213],"domain_scores_gemma":[0.9965102,0.001781648,0.0003152529,0.0005484434,0.0007696596,0.00007472945],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"simulation_or_modeling","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.001012356,0.0005632921,0.005457639,0.0003875622,0.0003256054,0.0003750047,0.000616021,0.4518682,0.02269287,0.03531002,0.02967101,0.4517205],"study_design_scores_gemma":[0.00001377403,0.0000232614,0.0004196837,0.000007274847,0.0000191557,0.00003918295,0.00001814435,0.9843626,0.002646513,0.01072535,0.001710303,0.00001474005],"study_design_candidate":"simulation_or_modeling","study_design_consensus":"simulation_or_modeling","genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.03263041,0.0003740277,0.9583014,0.0006533099,0.0001791953,0.0001559103,0.001331251,0.004400623,0.001973969],"genre_scores_gemma":[0.6015679,0.000570491,0.3655022,0.0006072427,0.0005885307,0.00108361,0.008636741,0.001281781,0.02016147],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.004715025,"threshold_uncertainty_score":0.01057285,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.0249597595725148,"score_gpt":0.2858597068773903,"score_spread":0.2608999473048755,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}