{"id":"W4409047962","doi":"10.1145/3721146.3721940","title":"Accelerating MoE Model Inference with Expert Sharding","year":2025,"lang":"en","type":"article","venue":"","topic":"Topic Modeling","field":"Computer Science","cited_by":2,"is_retracted":false,"has_abstract":true,"ca_institutions":"McGill University","funders":"","keywords":"Inference; Computer science; Artificial intelligence","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0009295943,0.0009936161,0.000918509,0.0005819701,0.0005264683,0.001238155,0.001892021,0.001015471,0.005310024],"category_scores_gemma":[0.005988672,0.000677569,0.0009366103,0.0006135271,0.0006825436,0.002856141,0.001838213,0.002062118,0.002110726],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0009516412,"about_ca_system_score_gemma":0.001922292,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.009820026,"about_ca_topic_score_gemma":0.02323771,"domain_scores_codex":[0.9993412,0.0001277475,0.00004184069,0.0002072627,0.0001739453,0.0001079904],"domain_scores_gemma":[0.9983432,0.0006791867,0.00009092583,0.000525909,0.0002386483,0.0001221796],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"simulation_or_modeling","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0007577512,0.0002082528,0.00473691,0.0002314386,0.0002541015,0.0003046797,0.0003760059,0.5573612,0.02445228,0.02060941,0.01881814,0.3718899],"study_design_scores_gemma":[0.00001189895,0.00001574456,0.00009863886,0.000003677367,0.000008432214,0.00002202463,0.00001586792,0.9887114,0.004195766,0.005990189,0.00091961,0.000006719914],"study_design_candidate":"simulation_or_modeling","study_design_consensus":"simulation_or_modeling","genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.03664187,0.0002613114,0.9427996,0.0002987502,0.00009821623,0.00005467193,0.0003400205,0.01734248,0.002163093],"genre_scores_gemma":[0.4661422,0.000166005,0.5244902,0.0004204661,0.00007721432,0.0001068165,0.001499857,0.001313886,0.005783461],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.009820026,"threshold_uncertainty_score":0.01952571,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.06416343562599425,"score_gpt":0.3062068170236703,"score_spread":0.2420433813976761,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}