ISCA Archive Odyssey 2026
ISCA Archive Odyssey 2026

Augmented State Space Speaker Clustering: Reformulating HMM Based Clustering To Improve Speaker Diarization

Anurag Chowdhury, Abhinav Misra, Yinong Wang, Bongjun Kim, Mark C. Fuhs, Monika Woszczyna

Speaker diarization systems segment a conversation into regions of speaker activity. While most segments contain speech from a single participant, a substantial portion includes multiple simultaneous speakers (multi-talker segments) or speech from non-participating speakers (background segments). Popular clustering-based approaches, such as VBx, assign these mixed segments to single-speaker clusters, effectively injecting noise into the speaker representations. We propose AS3C, a new HMM-based diarization framework that explicitly models multi-talker and background speech with dedicated HMM states and incorporates external speaker change-point information to guide state transitions. AS3C further supports optional integration of overlap speech detection (OSD) to generate overlap-aware outputs. AS3C significantly outperforms the baseline methods and achieves competitive or superior performance relative to the SOTA approaches on both the in-house DoPaCo and AMI datasets.