Speech encodes rich paralinguistic information such as demographics, voice quality, and health, yet existing audio foundation models offer limited zero-shot or out-of-domain (OOD) generalization to these tasks. We introduce SLAP (Speaker contrastive Language-Audio Pretraining), the first model aligning speech with natural language descriptions of diverse speaker and health metadata via contrastive learning. SLAP combines a Vision Transformer (ViT) audio encoder with a text encoder, trained on 3000 hours of audio across 8 datasets. We evaluate on 38 binary classification tasks spanning demographics, voice characteristics, and clinical assessments across 11 datasets in 6 languages. SLAP achieves 62.9% average F1 zero-shot, a 48% relative improvement over CLAP (42.4%), with strong OOD generalization to unseen clinical populations. With linear probing, SLAP reaches 69.3% F1 overall and best-in-class on health tasks (57.9% F1), surpassing larger foundation models including Whisper.