ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Domain Adaptation for Deepfake Audio Detection under Degraded Channel Conditions

Ayuto Tsutsumi, Akira Gotoh, Yuko Saito, Hiroki Matsuura, Sayaka Shiota

Recent zero-shot voice cloning methods generate speaker-faithful speech from short reference speech, posing a growing deepfake threat. Existing countermeasure (CM) models perform well on standard benchmarks, yet their robustness under degraded channel conditions has received little attention. In this paper, 11 zero-shot TTS/VC models are evaluated on speech recorded under three degraded conditions: phone calls, Telegram VoIP calls, and face-to-face interviews. State-of-the-art attacks raise the speaker verification EER from 0.91% to 11.63%, and pre-trained CMs fail under channel mismatch, with CM EERs approaching 50%. Domain adaptation using a mixture of original training data and in-domain degraded speech reduces the CM EER to below 2% for seen attacks. It markedly improves performance for unseen attacks, demonstrating its effectiveness under degraded channel conditions.