Back to News
Z.ai•December 11, 2025
GLM-TTS
Open Source
Audio
TTS
### TL;DR
GLM-TTS is a high-quality text-to-speech (TTS) synthesis system based on large language models, supporting zero-shot voice cloning and streaming inference. It utilizes a two-stage architecture combining a language model for speech token generation and a Flow Matching model for waveform synthesis. By introducing a Multi-Reward Reinforcement Learning framework, GLM-TTS significantly improves the expressiveness of generated speech, achieving more natural emotional control compared to traditional TTS systems.
Key Insights & Metrics
Pricing
Free
Cost structure
Version
1.0
Current release version
Hardware
CPU only
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- Zero-shot Voice Cloning: Clone any speaker's voice with just 3-10 seconds of prompt audio.
- RL-enhanced Emotion Control: Utilizes a multi-reward reinforcement learning framework (GRPO) to optimize prosody and emotion.
- High-quality Synthesis: Generates speech comparable to commercial systems with reduced Character Error Rate (CER).
- Phoneme-level Control: Supports "Hybrid Phoneme + Text" input for precise pronunciation control (e.g., polyphones).
- Streaming Inference: Supports real-time audio generation suitable for interactive applications.
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!