Back to News
Tencent Hunyuan, Shanghai Jiao Tong University, and Shanghai Innovation Institute•September 10, 2026
AuK
Open Source
Speech Generation
Voice Cloning
Audio Editing
Open Source
### TL;DR
AuK is a 1.5B-parameter foundational model designed to unify speech generation and editing through a single natural-language instruction interface. It supports a wide range of tasks including zero-shot TTS, content editing, speech enhancement, and source separation, utilizing a hybrid rectified-flow Transformer architecture.
Key Insights & Metrics
Pricing
Free (Open Source)
Cost structure
Version
N/A
Current release version
Hardware
An NVIDIA GPU with at least 8 GB VRAM (12 GB–16 GB+ recommended for the full base model and MLLM encoder) running on 16 GB+ system RAM and Python 3.10
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- Unified natural-language interface for generation, editing, enhancement, and separation
- 1.5B-parameter hybrid rectified-flow Transformer architecture
- AuK-Flash variant for 4-step high-speed inference
- Supports zero-shot and instruction-controlled speech generation
- Trained on 1.95 million hours of audio supervision
- MIT-licensed open-source weights and code
Ad

Master AI Marketing: Work Smarter, Not Harder
Unlock the power of AI to automate workflows and scale your results. Join 1.5M+ professionals and stay ahead of the curve with the latest AI insights.
Sponsored
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!