ISCA Archive SLTU 2018
ISCA Archive SLTU 2018

SVM Based Language Diarization for Code-Switched Bilingual Indian Speech Using Bottleneck Features

Spoorthy V, Veena Thenkanidiyoor, Dileep A. D

This paper proposes an SVM-based language diarizer for code switched bilingual Indian speech. Code-switching corresponds to usage of more than one language within a single utterance. Language diarization involves identifying code-switch points in an utterance and segmenting it into homogeneous language segments. This is very important for Indian context because every Indian is at least bilingual and code-switching is inevitable. For building an effective language diarizer, it is helpful to consider phonotactic features. In this work, we propose to consider bottleneck features for language diarization. Bottleneck features correspond to output of a narrow hidden layer of a multilayer neural network trained to perform phone state classification. The studies conducted using the standard datasets have shown the effectiveness of the proposed approach.