Overview
ChatTTS is an open-source text-to-speech model specifically optimized for conversational scenarios, such as dialogue tasks for large language model (LLM) assistants, conversational audio, and video introductions. Trained on approximately 100,000 hours of Chinese and English data, it delivers high-quality and natural speech synthesis. The model supports both Chinese and English, enabling multilingual applications.
Key Features
- Multi-language Support: Generates speech in Chinese and English.
- Large-Scale Training: Trained on 100,000 hours of diverse speech data for high naturalness.
- Dialogue Optimization: Designed for conversational tasks, providing fluid and natural interactions.
- Open Source: A base model trained on 40,000 hours is planned for release to enable community research and development.