ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Joint Timbral and Non-Timbral Speaker Anonymisation

Rayane Bakari, Olivier Le Blouch, Nicolas Gengembre, Nicholas Evans

Voice anonymisation aims to conceal speaker identity while preserving linguistic content. Most approaches target predominantly timbral cues and often overlook non-timbral cues such as prosody, rhythm, speaking style and accent, which may still leak speaker-specific information related to voice identity after anonymisation. With this paper, we propose a speaker anonymisation system that explicitly obfuscates both timbral and non-timbral cues. Extensive experiments conducted within the VoicePrivacy Challenge framework show improved protection against attacks exploiting non-timbral information compared to state-of-the-art systems. For evaluation, we use a pair of complementary automatic speaker verification models to demonstrate improved anonymisation robustness by 32% relative to attacks which target either timbral and non-timbral cues. Results also show stronger anonymisation comes at the cost of only moderate degradation to intelligibility and naturalness.