ISCA Archive Interspeech 2004
ISCA Archive Interspeech 2004

Keyword spotting for highly inflectional languages

Lubos Smidl, Ludek Müller

This paper presents our new keyword spotting system taking advantage of both the filler model and the confidence measure measure approaches. The novelty is in a non-standard connection of the filler and the keyword models together with introduction of a new confidence measure based on a keyword normalized score. In detail the paper deals with a decision block. Two methods are introduced. The first is based on comparison a keyword normalized score with a predefined decision threshold. The second uses three-layer feedforward neural network for decision if the keyword was or was not spoken. Results from the both presented methods are compared with the large vocabulary continuous speech recognition system used for keyword spotting. Obviously, LVCSR using a proper language model can give better results. Besides, it has higher CPU and memory demands. Furthermore, in many common situations the spontaneous language is mostly unconstrained and includes OOV words (keywords) such as names of peoples, companies, places, products etc., so the availability of an appropriate language model is very problematic.