StepFun AI

2025StepFun AI

Step-Audio-R1 v1.0

Open Source
LLMs
ML

Step-Audio-R1 is an advanced audio language model developed by StepFun AI, designed to enhance audio reasoning capabilities by grounding its reasoning in acoustic features. It introduces Modality-Grounded Reasoning Distillation (MGRD), an iterative training framework that shifts the model's reasoning from textual abstractions to acoustic properties, effectively addressing the 'inverted scaling' problem where performance degrades with longer reasoning. This model has demonstrated superior performance across various audio understanding and reasoning benchmarks, surpassing models like Gemini 2.5 Pro and achieving results comparable to Gemini 3 Pro.

Chain-of-Thought (CoT) Reasoning: Unlocks CoT reasoning capabilities in audio language models, generating reasoning chains grounded in acoustic features.
Modality-Grounded Reasoning Distillation (MGRD): An iterative training framework that shifts reasoning from textual abstractions to acoustic properties.
Superior Performance: Surpasses Gemini 2.5 Pro and is comparable to Gemini 3 across major audio reasoning tasks.
PricingFree
Version1.0
2026StepFun AI

Step 3.5 Flash v1.0

Open Source
LLM
LLMs

Step 3.5 Flash is an open-source foundation model engineered for advanced reasoning and agentic capabilities with exceptional efficiency. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token, achieving a generation throughput of 100–300 tokens per second. This design allows it to rival the reasoning depth of top-tier proprietary models while maintaining the agility required for real-time interaction.

Deep reasoning at speed with 3-way Multi-Token Prediction (MTP-3)
Robust engine for coding and agentic tasks with scalable reinforcement learning framework
Efficient long-context processing with 256K context window using Sliding Window Attention (SWA)
PricingFree
Version1.0
2026StepFun AI

STEP3-VL-10B v1.0

Open Source
Multimodal

STEP3-VL-10B is a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. Despite its compact 10B parameter footprint, STEP3-VL-10B excels in visual perception, complex reasoning, and human-centric alignment.

Unified pre-training on 1.2T multimodal tokens
Integration of a language-aligned Perception Encoder with a Qwen3-8B decoder
Scaled post-training pipeline with over 1,000 iterations of reinforcement learning
PricingFree
Version1.0
2025StepFun AI

Step-DeepResearch

Open Source
AI Agents
Agentic AI

Step-DeepResearch is a cost-effective, end-to-end deep research agent model designed for autonomous information exploration and professional report generation in open-ended research scenarios. It integrates atomic capabilities such as planning, information seeking, reflection, and report generation to perform comprehensive research tasks. The model is trained using a progressive pipeline that includes agentic mid-training, supervised fine-tuning, and reinforcement learning, enabling it to handle complex research workflows efficiently.

Atomic Capability Integration
Progressive Training Pipeline
Strong Performance Across Model Scales
RegionChina

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode