ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Subtract to Clean, Add to Enrich: Dual-Path Disentanglement for Speaker and Language Recognition

Aref Farhadipour

Self-supervised speech models capture rich acoustic and linguistic features, but adapting them to downstream tasks is computationally costly and prone to catastrophic forgetting. Moreover, cross-lingual Speaker Verification (ASV) and Language Identification (LID) suffer from an entanglement paradox: ASV degrades under cross-lingual mismatch, while LID exploits speaker identity as a shortcut. We propose a parameter-efficient dual-path architecture to address this issue. Using Wavelet Prefix Tuning on a W2V-BERT 2.0 backbone with MHFA heads, we introduce a Dual-Path Fusion module that disentangles representations via a subtractive path for language removal and an additive path for complementary feature retrieval. A bidirectional score penalization further reduces task-specific biases. Our method achieves 3.16% and 4.35% EER on ASV trials and 94.16% accuracy on LID benchmarks.