ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Harmonizing data augmentation and loss function for speaker recognition: examples with speed perturbation, mixup and mixout

Pierre-Michel Bousquet, Mickaël Rouvier

Data augmentation techniques have proved extremely effective across many fields of machine learning, including speaker recognition. While these techniques take the most of the available training data, they necessarily complicate the discrimination task during the learning phase, by forcing the model to take into account specific, unexpected or additional variability of the voice signal. In this paper, we show that modifying DNN components is essential for optimizing speaker recognition systems that include two data augmentation techniques: speed perturbation and mixup/mixout. We focus on adapting the optimization process of the DNN by reformulating the loss function that drives backpropagation and by designing a sampling algorithm, which selects the examples for batch parameter update of the network. The efficiency and robustness of our proposed approaches are tested on a wide range of evaluations, including test sets that span many types of mismatch.