Back to News
Tencent•January 27, 2026
Youtu-VL-4B-Instruct-GGUF
Open Source
CV
### TL;DR
Youtu-VL-4B-Instruct-GGUF is a lightweight yet robust Vision-Language Model (VLM) developed by Tencent. Built upon the Youtu-LLM with 4 billion parameters, it introduces the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm, enhancing visual perception and multimodal understanding. This model excels in both vision-centric and general multimodal tasks without the need for task-specific modules.
Key Insights & Metrics
Pricing
Free
Cost structure
Version
1.0
Current release version
Hardware
RTX 4090 GPU, 8GB RAM
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- Comprehensive Vision-Centric Capabilities
- Promising Performance with High Efficiency
- Unified Autoregressive Supervision
- Standard Architecture for Vision-Centric Tasks
- Versatile General-Purpose VLM
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!