{"id":"W2963381796","doi":"","title":"Text Categorization via Similarity Search - An Efficient and Effective Novel Algorithm.","year":2013,"lang":"en","type":"article","venue":"","topic":"Text and Document Classification Technologies","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Ottawa","funders":"","keywords":"Computer science; Similarity (geometry); Artificial intelligence; Centroid; Cluster analysis; Preprocessor; Outlier; Metric (unit); Categorization; Text categorization; Point (geometry); Class (philosophy); Similarity measure; Pattern recognition (psychology); Algorithm; Mathematics; Image (mathematics)","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.001154257,0.001005275,0.001606086,0.005112691,0.0009622465,0.00166594,0.002094193,0.001661175,0.003703929],"category_scores_gemma":[0.004211647,0.0004100468,0.001006333,0.005742155,0.0006847068,0.003933484,0.001796985,0.001254973,0.005335423],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.000782432,"about_ca_system_score_gemma":0.001652709,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.002122209,"about_ca_topic_score_gemma":0.003090663,"domain_scores_codex":[0.9977671,0.0003875529,0.0002029285,0.0005165889,0.001017873,0.0001079249],"domain_scores_gemma":[0.9987207,0.0003717825,0.000128999,0.0002816681,0.0004372186,0.00005954551],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"bench_or_experimental","study_design_scores_codex":[0.0001510674,0.0002963168,0.00117113,0.0002183753,0.0001044038,0.0001020068,0.0001106463,0.007291809,0.01206886,0.007219689,0.01950626,0.9517595],"study_design_scores_gemma":[0.000219602,0.0003768452,0.002760408,0.00009331626,0.0001148346,0.001679006,0.0003487896,0.8460934,0.02602074,0.06087798,0.06132947,0.00008564097],"study_design_candidate":"bench_or_experimental","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.009021591,0.00154571,0.9808567,0.0004322215,0.0002446361,0.0004397533,0.0005328787,0.004276154,0.002650415],"genre_scores_gemma":[0.06015883,0.0004734429,0.9304464,0.0002332256,0.0001797242,0.0003985765,0.002364887,0.0001720536,0.005572775],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.005112691,"threshold_uncertainty_score":0.01239085,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.01306940863528285,"score_gpt":0.2493323822246399,"score_spread":0.2362629735893571,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}