{"id":"W3158310581","doi":"10.1111/rssc.12483","title":"Clustering and Automatic Labelling Within Time Series of Categorical Observations—With an Application to Marine Log Messages","year":2021,"lang":"en","type":"article","venue":"Journal of the Royal Statistical Society Series C (Applied Statistics)","topic":"Time Series Analysis and Forecasting","field":"Computer Science","cited_by":3,"is_retracted":false,"has_abstract":true,"ca_institutions":"","funders":"NRC Steacie Institute for Molecular Sciences","keywords":"Computer science; Cluster analysis; Data mining; Set (abstract data type); A priori and a posteriori; Categorical variable; Inference; Bayesian probability; Series (stratigraphy); State (computer science); Hierarchical clustering; Identification (biology); Algorithm; Pattern recognition (psychology); Artificial intelligence; Machine learning","routes":{"ca_aff":false,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":true},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.002345048,0.0004512354,0.0007579435,0.002812926,0.0007921253,0.001148973,0.001576678,0.001260336,0.001080715],"category_scores_gemma":[0.01348777,0.000311937,0.0005653261,0.002602661,0.0006930814,0.00100523,0.001045452,0.00125543,0.0006106063],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.001074196,"about_ca_system_score_gemma":0.0009081833,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.01280371,"about_ca_topic_score_gemma":0.0135469,"domain_scores_codex":[0.9986027,0.0005894562,0.00009046176,0.0003753568,0.0002364266,0.0001056077],"domain_scores_gemma":[0.9896039,0.007093653,0.0008056727,0.001044531,0.001224282,0.0002280624],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0006602831,0.0004948401,0.02397577,0.0002884559,0.0001429022,0.0003756325,0.001275903,0.3781354,0.01048298,0.01988052,0.006549467,0.5577378],"study_design_scores_gemma":[0.000009009043,0.00001906131,0.002413015,0.000008665209,0.000005403604,0.00002410726,0.00006512396,0.9883415,0.0008796931,0.007611279,0.0006075418,0.00001551341],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.1206421,0.0002288257,0.8750854,0.0003980666,0.00006416943,0.0001415568,0.0006424989,0.002024652,0.0007727606],"genre_scores_gemma":[0.5534413,0.0001274576,0.4429001,0.00007508675,0.0000937325,0.0001512857,0.001619478,0.0001757964,0.001415655],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.01280371,"threshold_uncertainty_score":0.02545834,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.01078433395384456,"score_gpt":0.2181222175034086,"score_spread":0.2073378835495641,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}