vLLM v0.25.0
Open Source
LLM
Inference Engine
vLLM is a high-throughput, memory-efficient inference and serving engine for large language models. The v0.25.0 release introduces Model Runner V2 as the default engine for dense models, removes legacy PagedAttention, and adds extensive support for new models, hardware, and performance optimizations. Key technical advancements include improved support for heterogeneous speculative decoding, a new streaming parser engine, and enhanced support for a wide array of architectures such as GLM-5, MiniMax-M3, and various multimodal models.
Model Runner V2 as the default engine for dense models
Universal speculative decoding for heterogeneous vocabularies
New Streaming Parser Engine for tool-call and reasoning frameworks
PricingOpen Source
Version0.25.0