ISCA Archive SpeechProsody 2026
ISCA Archive SpeechProsody 2026

Detection of Lombard Speech Using Different Model Architectures and Speech Features

Judith Bauer, Frank Zalkow, Meinard Müller, Christian Dittmar

The Lombard effect refers to the speech modulations that occur when people talk in background noise, such as higher pitch and loudness, slower or clearer articulation, and shifts in vocal timbre. It is usually studied by recording speakers who hear controlled noise through headphones, producing clean speech signals that still carry the acoustic imprint of this noise-induced adaptation. Since stronger noise typically leads to stronger speech modulations, we frame Lombard detection as a regression task: predicting the noise level from clean speech as a measure of the strength of the Lombard effect. Using publicly available datasets, we train and compare three model families: convolutional detectors, convolutional detectors with an LSTM layer, and transformer-based detectors. As a main contribution, we examine how different feature choices affect performance by combining one rich speech representation (phonetic posteriorgrams, mel spectrograms, or WavLM features) with a set of interpretable acoustic cues, including pitch, energy, phoneme duration, first formant, and spectral tilt. Evaluation on a held-out dataset shows that convolutional models generalize most reliably and especially perform well when trained with WavLM features. Across the rich speech representations, convolutional models benefit especially from pitch and spectral tilt information, resulting in statistically significantly lower loss values.