Back to News
inclusionAI•February 11, 2026
Ming-flash-omni 2.0
Open Source
Audio
TTS
### TL;DR
Ming-flash-omni 2.0 is an open-source, state-of-the-art multimodal large language model developed by inclusionAI. It leverages the Ling-2.0 architecture, a Mixture-of-Experts (MoE) framework comprising 100 billion total parameters, with 6 billion active parameters per token. This design enables efficient scaling and empowers unified multimodal intelligence across vision, speech, and language, representing a significant advancement toward Artificial General Intelligence (AGI).
Key Insights & Metrics
Pricing
Free
Cost structure
Version
2.0
Current release version
Hardware
NVIDIA A100 GPU or equivalent, 32GB RAM, CUDA-enabled system
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region
Key Features
- Expert-level Multimodal Cognition: Accurately identifies plants, animals, cultural references, and artifacts, delivering expert-level analysis.
- Immersive and Controllable Unified Acoustic Synthesis: Integrates speech, audio, and music within a single channel, enabling zero-shot voice cloning and nuanced attribute control.
- High-Dynamic Controllable Image Generation and Manipulation: Unifies segmentation, generation, and editing, allowing for sophisticated spatiotemporal semantic decoupling.
- State-of-the-Art Performance: Achieves new benchmarks in contextual ASR, dialect-aware ASR, text-to-image generation, and generative segmentation.
- Open-Source Accessibility: Released under the MIT license, promoting transparency and community collaboration.
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!