ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Flow-Enhanced Language Embeddings for Robust Language Recognition

Tianyu Cao, Laureano Moro Velazquez, Jesus Villalba, Thomas Thebaud, Najim Dehak

Language identification (LID) systems are widely used in speech analytics, multilingual routing, and spoken-language interfaces, but their performance often drops in real-world conditions because of domain mismatch, including background noise, reverberation, channel effects, codecs, and far-field recording. We propose an embedding-level generative enhancement method based on flow matching that refines language embeddings from a pre-trained LID model. During training, we build paired embeddings from clean utterances and their distorted versions, and learn a continuous-time transformation that maps corrupted embeddings toward clean ones while preserving language-discriminative structure. At inference, any embedding is passed through the learned flow to obtain a refined, more robust representation. The method requires no changes to the LID pipeline and no language labels, while improving LID accuracy under challenging conditions without hurting standard performance.