Back to News
Qwen•January 8, 2025
Qwen3-VL-Embedding
Open Source
Multimodal
### TL;DR
Qwen3-VL-Embedding is a series of models designed for high-precision multimodal retrieval and ranking. Built upon the Qwen3-VL foundation model, it processes diverse inputs—including text, images, screenshots, and videos—by mapping them into a unified representation space. This enables efficient similarity computation and retrieval across different modalities.
Key Insights & Metrics
Pricing
Open source; free to use under the Apache-2.0 license
Cost structure
Version
2B and 8B parameter sizes
Current release version
Hardware
Specific hardware requirements are not specified; performance may vary based on model size and deployment environment
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- Supports over 30 languages
- Processes text, images, screenshots, and videos
- Offers flexible embedding dimensions (64–2048)
- Utilizes Matryoshka Representation Learning for customizable embedding sizes
- Handles inputs up to 32k tokens
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!