← all releases
OrcaRouter·Sep 26, 2026·v2·open source
OrcaSAQ-2 27B
OrcaSAQ-2 27B is a highly compressed, sensitivity-aware mixed-precision quantized version of the Qwen3.8-27B model. It reduces the storage footprint by over 77% while maintaining high fidelity to the original model, making it suitable for deployment on hardware with limited GPU memory.
quantizationreasoningagenticvllm
overview
OrcaSAQ-2 27B is a highly compressed, sensitivity-aware mixed-precision quantized version of the Qwen3.8-27B model. It reduces the storage footprint by over 77% while maintaining high fidelity to the original model, making it suitable for deployment on hardware with limited GPU memory.
key features
- 013-bit sensitivity-aware mixed-precision quantization
- 02Supports MTP speculative decoding and tool calling
- 03Optimized for long-horizon agents and complex reasoning
- 04262K context window support
- 05Compatible with vLLM for production serving
related productsbrowse all →