This paper describes the process of building HMM-based speech synthesis system (HTS) voices for our participation in the Blizzard Challenge 2009. Out of the two languages required (English and Mandarin Chinese) we only built three Mandarin Chinese voices for main hub (MH) and two spoke (MS1 and MS2) tasks. According to the evaluation results, our MH voice got 3 points for both mean opinion scores (MOS) and similarity tests. Beside, 12.2% and 17% pinyin error rates (without (PER) and with tone (PTER), respectively) and 23% character error rate (CER) were achieved for intelligibility test. Moreover, our MS2 voice achieved 4 and 3 points for MOS and similarity test, respectively. In conclusion, we now have reasonable text-to-speech (TTS) baselines (at least for Mandarin Chinese) for developing our own advanced prosody model in the future.