ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision

Angelika Andò, Auguste Crabeil, Quentin Spinat, Adrien Lesage, Rachid Riad

Speech encodes rich paralinguistic information such as demographics, voice quality, and health, yet existing audio foundation models offer limited zero-shot or out-of-domain (OOD) generalization to these tasks. We introduce SLAP (Speaker contrastive Language-Audio Pretraining), the first model aligning speech with natural language descriptions of diverse speaker and health metadata via contrastive learning. SLAP combines a Vision Transformer (ViT) audio encoder with a text encoder, trained on 3000 hours of audio across 8 datasets. We evaluate on 38 binary classification tasks spanning demographics, voice characteristics, and clinical assessments across 11 datasets in 6 languages. SLAP achieves 62.9% average F1 zero-shot, a 48% relative improvement over CLAP (42.4%), with strong OOD generalization to unseen clinical populations. With linear probing, SLAP reaches 69.3% F1 overall and best-in-class on health tasks (57.9% F1), surpassing larger foundation models including Whisper.