{"id":"W4414260593","doi":"10.1145/3767334","title":"SuperBench: A Proactive Validation System for Improving Reliability of Cloud AI Infrastructure","year":2025,"lang":"en","type":"article","venue":"ACM Transactions on Computer Systems","topic":"Software System Performance and Reliability","field":"Computer Science","cited_by":1,"is_retracted":false,"has_abstract":true,"ca_institutions":"Microsoft (Canada)","funders":"","keywords":"Testbed; Cloud computing; Benchmark (surveying); Validator; Reliability (semiconductor); Fault tolerance; Root cause","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.005290049,0.001926637,0.0006529811,0.001815603,0.0007341967,0.001556367,0.003537863,0.0008648678,0.004670304],"category_scores_gemma":[0.01561269,0.0008909297,0.0007073433,0.0006548261,0.00108821,0.003416703,0.002454962,0.001772173,0.001945298],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.001030315,"about_ca_system_score_gemma":0.002557053,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.003501555,"about_ca_topic_score_gemma":0.004276811,"domain_scores_codex":[0.996749,0.0009310332,0.000262986,0.000529863,0.001206667,0.0003204451],"domain_scores_gemma":[0.9872904,0.003629618,0.001283737,0.004044573,0.003227134,0.0005244917],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.003324983,0.001068118,0.04591211,0.001733636,0.0005864867,0.0009637452,0.001345876,0.1728772,0.1987352,0.01212591,0.1332856,0.4280413],"study_design_scores_gemma":[0.0002881005,0.001233587,0.008857175,0.0001449471,0.0001490537,0.0003319319,0.0001748241,0.8074752,0.1384528,0.006207909,0.03648534,0.0001990683],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.1863808,0.001611641,0.4407291,0.0007953769,0.0005333485,0.0009851011,0.003368421,0.3548548,0.01074142],"genre_scores_gemma":[0.7559081,0.0005137501,0.2180717,0.0008518064,0.0001300275,0.0005615971,0.007190829,0.01011762,0.006654531],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.005290049,"threshold_uncertainty_score":0.02797675,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.008737100156824773,"score_gpt":0.2430366704539447,"score_spread":0.2342995702971199,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}