Back to News
Tencent Hunyuan, Shanghai Jiao Tong University, and Shanghai Innovation InstituteSeptember 10, 2026

AuK

Open Source
Speech Generation
Voice Cloning
Audio Editing
Open Source

Explore AuK

Visit the official website to learn more and get started

### TL;DR

AuK is a 1.5B-parameter foundational model designed to unify speech generation and editing through a single natural-language instruction interface. It supports a wide range of tasks including zero-shot TTS, content editing, speech enhancement, and source separation, utilizing a hybrid rectified-flow Transformer architecture.

Key Insights & Metrics

Pricing
Free (Open Source)
Cost structure
Version
N/A
Current release version
Hardware
An NVIDIA GPU with at least 8 GB VRAM (12 GB–16 GB+ recommended for the full base model and MLLM encoder) running on 16 GB+ system RAM and Python 3.10
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region

Key Features

  • Unified natural-language interface for generation, editing, enhancement, and separation
  • 1.5B-parameter hybrid rectified-flow Transformer architecture
  • AuK-Flash variant for 4-step high-speed inference
  • Supports zero-shot and instruction-controlled speech generation
  • Trained on 1.95 million hours of audio supervision
  • MIT-licensed open-source weights and code
Ad
Master AI Marketing: Work Smarter, Not Harder

Master AI Marketing: Work Smarter, Not Harder

Unlock the power of AI to automate workflows and scale your results. Join 1.5M+ professionals and stay ahead of the curve with the latest AI insights.

Sponsored

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!