Back to News
NVIDIAJanuary 19, 2026

Qwen3-8B-DMS-8x

Open Source
ML

Explore Qwen3-8B-DMS-8x

Visit the official website to learn more and get started

### TL;DR

Qwen3-8B-DMS-8x is a derivative of Qwen3-8B that integrates Dynamic Memory Sparsification (DMS) with an 8x compression ratio during inference. DMS adaptively sparsifies the key-value (KV) cache to reduce memory footprint and improve throughput and latency for long-context and reasoning generations. The method learns per-head eviction policies that interpolate between a sliding window over the last 512 tokens and full attention. Inference-time code is provided with the checkpoint.

Key Insights & Metrics

Pricing
Non-commercial research and educational use only under the NVIDIA License
Cost structure
Version
Qwen3-8B-DMS-8x
Current release version
Hardware
NVIDIA GPU-accelerated systems; compatible with NVIDIA Ampere, Blackwell, and Hopper architectures; preferred operating system is Linux
Compute requirements
Category
Open Source
Licensing model
Region
United Kingdom
Primary region

Key Features

  • Integrates Dynamic Memory Sparsification (DMS) with 8x compression ratio during inference
  • Reduces memory footprint and improves throughput and latency for long-context and reasoning generations
  • Provides inference-time code with the checkpoint

Related Releases

SkyRL tx

SkyRL tx is an open-source library that implements a backend for the Tinker API, enabling users to set up their own Tinker-like services on personal hardware. It supports end-to-end reinforcement learning (RL) and offers significantly faster sampling. The library is designed to be modular, allowing easy prototyping of new training algorithms, environments, and execution plans without compromising usability or speed.

NovaSky AINov 3
Open

Step-Audio-R1

Step-Audio-R1 is an advanced audio language model developed by StepFun AI, designed to enhance audio reasoning capabilities by grounding its reasoning in acoustic features. It introduces Modality-Grounded Reasoning Distillation (MGRD), an iterative training framework that shifts the model's reasoning from textual abstractions to acoustic properties, effectively addressing the 'inverted scaling' problem where performance degrades with longer reasoning. This model has demonstrated superior performance across various audio understanding and reasoning benchmarks, surpassing models like Gemini 2.5 Pro and achieving results comparable to Gemini 3 Pro.

StepFun AINov 29
Open

SINQ

SINQ (Sinkhorn-Normalized Quantization) is a novel, fast, and high-quality quantization method designed to make any Large Language Model (LLM) smaller while preserving accuracy. It offers a plug-and-play, model-agnostic technique that delivers state-of-the-art performance for LLMs without sacrificing accuracy.

HuaweiNov 14
Open

AIBuildAI

AIBuildAI is an AI agent that autonomously constructs AI models. Given a specific task, it initiates an agent loop to analyze the problem, design models, and execute training processes, all without human intervention.

AIBuildAI
Open

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode