This paper presents two letter-to-sound (LTS) methods in building a small-footprint multilingual text-to-speech (TTS) engine. For the languages where there exist a systematic relationship between a word format and its pronunciation, we employ a rule-based method. Otherwise, we use a training-based method. In the second method, we adopted optimal sequence to implement the process of letter-to-phoneme alignment, and use CART to train the decision tree and store the results in an efficient way. Despite their merits and disadvantages, the experimental results on six languages demonstrate these methods are effective to reach an acceptable LTS precision within a reasonable small size.