← all releases
OrcaRouter·Sep 26, 2026·v2·open source

OrcaSAQ-2 27B

OrcaSAQ-2 27B is a highly compressed, sensitivity-aware mixed-precision quantized version of the Qwen3.8-27B model. It reduces the storage footprint by over 77% while maintaining high fidelity to the original model, making it suitable for deployment on hardware with limited GPU memory.

quantizationreasoningagenticvllm
visit website hugging face ↗
overview

OrcaSAQ-2 27B is a highly compressed, sensitivity-aware mixed-precision quantized version of the Qwen3.8-27B model. It reduces the storage footprint by over 77% while maintaining high fidelity to the original model, making it suitable for deployment on hardware with limited GPU memory.

key features
  • 013-bit sensitivity-aware mixed-precision quantization
  • 02Supports MTP speculative decoding and tool calling
  • 03Optimized for long-horizon agents and complex reasoning
  • 04262K context window support
  • 05Compatible with vLLM for production serving
related productsbrowse all →