{"id":"W4416143831","doi":"10.65148/ecn/2025019","title":"Personalized Text to Speech Synthesis through Few Shot Speaker Adaptation with Contrastive Learning","year":2025,"lang":"en","type":"article","venue":"Elaris Computing Nexus","topic":"Speech Recognition and Synthesis","field":"Computer Science","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"Trinity College","funders":"","keywords":"Naturalness; Similarity (geometry); Speaker recognition; Mean opinion score; Speech synthesis; Encoder; Feature learning; Word error rate; Speaker diarisation","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0004100564,0.0006769203,0.0005230203,0.000251058,0.0001777767,0.0003701316,0.0007504963,0.0006083403,0.002355356],"category_scores_gemma":[0.00119888,0.0002073735,0.0006826473,0.0001907444,0.000297825,0.0007980988,0.0007719381,0.0009426766,0.001294953],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0002252902,"about_ca_system_score_gemma":0.0003073092,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.0009662585,"about_ca_topic_score_gemma":0.001760877,"domain_scores_codex":[0.9997205,0.00005665706,0.00001444574,0.0001130095,0.00007430035,0.00002118495],"domain_scores_gemma":[0.9996647,0.0001609155,0.00002406263,0.00006057457,0.00006644332,0.00002328584],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0007681867,0.0002960474,0.0008727784,0.0002031753,0.0001504898,0.0002932292,0.0002170679,0.1797822,0.2545539,0.003256426,0.003536386,0.5560701],"study_design_scores_gemma":[0.00003537472,0.0001866112,0.0005192179,0.000008539743,0.00003267374,0.0001564227,0.00003118356,0.9420244,0.05254997,0.002340989,0.002089776,0.00002487674],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.04360974,0.0003538465,0.95006,0.00009109481,0.0001255706,0.00006738661,0.0001263055,0.003400607,0.002165473],"genre_scores_gemma":[0.6339163,0.0002332816,0.3573038,0.0002594212,0.00009064517,0.0002073882,0.00082047,0.0005184396,0.006650294],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.002355356,"threshold_uncertainty_score":0.007879496,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.02728395361844038,"score_gpt":0.2702029549770197,"score_spread":0.2429190013585793,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}