བོད་སྐད་ནི་ང་ཚོའི་རིག་གཞུང་དང་ངོ་བོའི་ཆ་ཤས་གལ་ཆེན་ཞིག་ཡིན།

དེང་རབས་ཀྱི་ལག་རྩལ་གསར་པས་བོད་སྐད་ལ་གོ་སྐབས་གསར་པ་སྐྲུན་ཐུབ།

Introducing Monlam TTS: Giving Tibetan a Digital Voice

Monlam TTS is the state-of-the-art Tibetan speech generation model developed by the Monlam AI team, designed to transform written Tibetan text into natural and expressive speech.

Monlam TTS

Introducing Monlam TTS

Humans communicate primarily through speech. However, Tibetan language still lacks many of the high-quality speech technologies available for widely supported languages. Building AI systems that can fully interact in Tibetan requires two key abilities: understanding speech and generating speech. With Monlam STT enabling AI systems to understand Tibetan speech, Monlam TTS completes the other half of the interaction by enabling AI to speak Tibetan. Together, these technologies represent an important step towards creating complete Tibetan voice-based AI systems.

We developed Monlam TTS to transform Tibetan text into natural speech, making Tibetan digital content more accessible for education, communication, and future AI applications. Monlam Text-to-Speech is a Tibetan speech generation model designed to convert Tibetan text into natural sounding speech. The system is built specifically for performance in the Tibetan language and can produce realistic speech in over 9+ voices with dialect variations. Equipped with a robust pipeline, the model can gracefully handle even long texts without any inconsistencies. The voice generated is natural-sounding with accurate pronunciations and smooth rhythm similar to human speech.

གལ་སྲིད་དུས་ཚོད་ཡོད་ན། ང་ཚོ་མཉམ་དུ་ཇ་ཁང་ལ་ཕྱིན་ནས་ལས་ཀའི་སྐོར་གོ་བསྡུར་ཞིག་བྱས་ན་བསམ་གྱིན་འདུག ཧ་ཧ་ཧ་ ཁྱེད་རང་ག་དུས་ལྷོད་ལྷོད་ཆགས་ན་ང་ལ་ལན་ཞིག་གནང་རོགས།

Why Tibetan Text-to-Speech is challenging

Building a high-quality Tibetan text-to-speech system inherently gets accompanied by many unique challenges. Unlike widely supported languages, Tibetan has limited publicly available speech resources for training modern AI models. In addition, variations across Tibetan dialects make accurate pronunciation and natural speech generation challenging.

Another major challenge is generating speech that goes beyond correct pronunciation and captures rhythm, pacing, and natural expression to create voices that feel comfortable for people to listen to. Finally, there is a highly limited amount of previous work done on this that can be used for reference and research purposes.

To address these challenges, we created comprehensive pipelines for the corpus collection and creation of the training data. The latest and best techniques available were adapted for compatibility with the Tibetan speech generation and developed a custom tokenizer for Tibetan text.

Monlam TTS text

Building the Model

High quality AI systems are built upon high quality data. For the development of Monlam TTS, we trained the model using over 2,400 hours of carefully prepared Tibetan speech data. For a low resource language such as Tibetan, creating a dataset of this scale required extensive collection, processing, and quality verification by dedicated teams.

Experience Monlam TTS

Amdo - Female 1

Lhasa - Male 1

Lhasa - Female 1

བཀྲ་ཤིས་བདེ་ལེགས། ང་ད་ལྟ་ལས་ཁུངས་ལ་འབྱོར་སོང་། ཁྱེད་རང་དེ་རིང་ལས་ཀ་བྲེལ་བ་ཚ་པོ་ཡོད་པེ།

Monlam TTS typography

Applications and future vision

Monlam TTS’s ability to generate natural Tibetan speech opens the door to a new dimension of accessibility to the language. People can now experience Tibetan digital content not only through reading, but also through listening. This opens many new possibilities ranging from creating audiobooks of Tibetan literature, to personal AI assistants with speech to speech capabilities.

We are actively working on the next steps for this latest development. The model will be available for use to everyone as an API through our new (yet to be released) Melong AI Studio along with other models. We hope that this enables its integration in independent projects working for a positive impact on Tibetan language and culture preservation. We are refining the model further to soon have improved emotional expression in the generated speech and add many more new regional dialects as new voices.

The current capability of the Monlam TTS represents a major step in our goal of digital inclusion and preservation of the Tibetan language through AI. Together with the Monlam STT model, this lets us build a complete interactive Tibetan speech system and lays the groundwork for the future of the language in AI and new technologies. We hope that this technology can help everyone access the Tibetan language, be it people newly learning Tibetan or people who already know the language and want to use it in new digital environments.