ISCA Archive Interspeech 2004
ISCA Archive Interspeech 2004

Using machine learning to cope with imbalanced classes in natural speech: evidence from sentence boundary and disfluency detection

Yang Liu, Elizabeth Shriberg, Andreas Stolcke, Mary Harper

We investigate machine learning techniques for coping with highly skewed class distributions in two spontaneous speech processing tasks. Both tasks, sentence boundary and disfluency detection, provide important structural information for downstream language processing modules. We examine the effect of data set size, task, sampling method (no sampling, downsampling, oversampling, and ensemble sampling), and learning method (bagging, ensemble bagging, and boosting) for a decision tree prosody model.