Huawei
2025•Huawei
SINQ v1.0
Open Source
Infrastructure
LLMs
SINQ (Sinkhorn-Normalized Quantization) is a novel, fast, and high-quality quantization method designed to make any Large Language Model (LLM) smaller while preserving accuracy. It offers a plug-and-play, model-agnostic technique that delivers state-of-the-art performance for LLMs without sacrificing accuracy.
Calibration-free quantization
Supports both symmetric and asymmetric quantization
NF4 support
PricingFree
Version1.0
2026•Huawei
Huawei 950PR Inference Chip v950PR
Paid
Infrastructure
ML
Huawei released the 950PR Inference Chip, targeting developers who need to deploy high-speed AI inference locally at the edge. Designed to rival NVIDIA's mid-range inference cards, the 950PR offers a competitive alternative for edge computing deployments without relying on cloud infrastructure.
High-speed local AI inference for edge computing
Rivals NVIDIA mid-range inference cards on performance
Targeted at on-premise and edge deployment scenarios
Version950PR
RegionCN