Aggressive Data Scaling
The team is continuing to expand both classical and modern Tibetan corpora while improving preprocessing and synthetic generation systems. Future training phases aim to scale data volume to approximately 3× the current corpus size.