ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

From Single-Channel Foundations to Multi-Speaker and Multi-Modal Understanding

Lukáš Burget

Recent advances in foundation models such as Whisper and WavLM have transformed automatic speech recognition, yet most systems still assume a single, clean speaker. This talk traces the progression toward models that can process and understand natural multi-speaker conversations. I will discuss how large pre-trained speech models can be extended to multi-channel input and spatially aware processing, how speaker diarization and recognition can be unified within a single framework, and how efficient model compression enables real-time deployment. Together, these developments move the field from modular, error-prone pipelines toward integrated systems capable of attributing and transcribing overlapping speech in realistic acoustic conditions. Looking ahead, I will outline ongoing efforts to extend these ideas beyond audio, toward audio-visual modeling and toward combining speech encoders with large language model decoders that can summarize or interpret conversations. These directions reflect a broader goal in speech technology: bringing machines closer to understanding who is speaking, what is being said, and ultimately what it means.