Back to News
NVIDIA•August 11, 2026
Nemotron 3.5 Lightning
Open Source
text-generation
mixture-of-experts
mamba
nvidia
### TL;DR
NVIDIA Nemotron 3.5 Lightning 30B-A3B-NVFP4 is a high-performance, latency-optimized large language model featuring a hybrid Mamba-2 and Mixture-of-Experts (MoE) architecture. Designed for efficient agentic workflows, it utilizes 30B total parameters with 3B active parameters per token to deliver fast, accurate execution for specialized tasks.
Key Insights & Metrics
Pricing
Free (Open Weights)
Cost structure
Version
3.5
Current release version
Hardware
NVIDIA GPU-accelerated systems (e.g., H100, DGX Spark, or Ampere-based GPUs for W4A16 quantization)
Compute requirements
Category
Open Source
Licensing model
Region
United States
Primary region
Key Features
- Hybrid Mamba-2 and Mixture-of-Experts (MoE) architecture
- 30B total parameters with 3B active parameters per token
- 1 million token context window
- Optimized for low-latency, high-throughput agentic tasks
- Native support for tool calling and reasoning control
- Commercial-ready under OpenMDW-1.1 license
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!