ISCA Archive Interspeech 2004
ISCA Archive Interspeech 2004

The use of typical sequences for robust speaker identification

Mohamed Mihoubi, Douglas O'Shaughnessy, Pierre Dumouchel

Speech field is affected by accidental structures such as spurious events and artifacts (breath mouth, lip clicks etc). Because the whole utterance is used during training and the identification process, these factors may represent confusable acoustic classes which do not contribute to the performance of the speaker recognition system. In this paper we propose a new approach for extracting as much essential information as possible from the acoustic data in order to estimate more robust speaker models. To this end, we make use of the Asymptotic Equipartition Property (AEP) originated from the information theory to generate a reduced feature subspace termed typical set. Results on the Spidre corpus show that the method leads to an appreciable improvement (averaging 10 % gain in performance) over the baseline system.