Inception Labs
2026•Inception Labs
Mercury 2 v2.0
Paid
LLMs
Mercury 2 is a cutting-edge diffusion-based large language model (dLLM) developed by Inception Labs, designed to deliver ultra-fast and efficient AI-driven text generation. By leveraging diffusion technology, Mercury 2 generates multiple tokens simultaneously, achieving speeds over 1,000 tokens per second on NVIDIA Blackwell GPUs, significantly outperforming traditional autoregressive models.
Diffusion-based architecture for parallel token generation
Over 1,000 tokens per second processing speed on NVIDIA Blackwell GPUs
Competitive performance with leading speed-optimized models
Version2.0
RegionUnited States
2026•Inception Labs
Mercury 2 vMercury 2
Paid
LLM
Diffusion Model
Mercury 2 is the world's fastest reasoning LLM, built on Inception Labs' diffusion architecture. It generates tokens in parallel rather than sequentially, achieving over 1,000 tokens/sec on NVIDIA Blackwell GPUs — making reasoning-grade quality viable within real-time latency budgets for agents, voice, and search pipelines.
1,009 tokens/sec on NVIDIA Blackwell — parallel diffusion decoding delivers >5x faster generation than autoregressive models at the same quality tier
Tunable reasoning with 128K context, native tool use, and schema-aligned JSON output — production-ready for agentic loops, coding, voice, and RAG pipelines
OpenAI API-compatible — drop in as a replacement with no rewrites required; $0.25/M input · $0.75/M output
Pricing$0.25/M input · $0.75/M output tokens
VersionMercury 2