ISCA Archive SSW 2025
ISCA Archive SSW 2025

Exploring Language Dependency in Ultrasound-to-Speech Synthesis

Ibrahim Ibrahimov, Csaba Zainkó, Gábor Gosztolya

Articulation-to-speech synthesis using ultrasound tongue imaging is a promising approach for Silent Speech Interfaces. However, its effectiveness is hindered by challenges such as session and speaker dependency, dataset scarcity and language variability. This study explores the language dependency of an ultrasound-to-speech synthesis system, consisting of a 2D-CNN to map ultrasound tongue images to mel spectrograms and a HiFi-GAN vocoder. The CNNs were trained on Azerbaijani recordings collected from three native speakers, each recorded in a single session containing both Azerbaijani (L1) and English (L2) sentences, and were then used to generate mel spectrograms for both languages. While the CNNs showed language dependency with lower mean squared error on L1, the mel-cepstral distortion of the synthesized speech did not reflect this, revealing the language bias of the vocoder. These results demonstrate the importance of considering language-specific factors in silent speech synthesis.