The present research focuses on the acquisition and annotation of vocal resources for emotion detection. We are interested in detecting emotions occurring in abnormal situations and particularly in detecting "fear". The present study considers a preliminary database of audiovisual sequences extracted from movie fictions. The sequences selected provide various manifestations of target emotions and are described with a multimodal annotation tool. We focus on audio cues in the annotation strategy and we use the video as support for validating the audio labels. The present article deals with the description of the methodology of data acquisition and annotation. The validation of annotation is realized via two perceptual paradigms in which the +/-video condition in stimuli presentation varies. We show the perceptual significance of the audio cues and the presence of target emotions.