Back to News
openbmbFebruary 10, 2026

UltraData-Math

Open Source
ML

Explore UltraData-Math

Visit the official website to learn more and get started

### TL;DR

UltraData-Math is a large-scale, high-quality mathematical pre-training dataset designed to enhance mathematical reasoning in large language models (LLMs). It comprises over 290 billion tokens across three progressive tiers: L1 (170.5B tokens from web corpus), L2 (33.7B tokens from quality-selected data), and L3 (88B tokens from multi-format refined data). This dataset has been utilized in the mathematical pre-training of the MiniCPM Series models.

Key Insights & Metrics

Pricing
Free
Cost structure
Version
1.0
Current release version
Hardware
CPU only
Compute requirements
Category
Open Source
Licensing model
Region
China
Primary region

Key Features

  • Comprehensive dataset with over 290 billion tokens
  • Three progressive tiers (L1, L2, L3) for systematic enhancement
  • Applied in pre-training MiniCPM Series models

Related Releases

SkyRL tx

SkyRL tx is an open-source library that implements a backend for the Tinker API, enabling users to set up their own Tinker-like services on personal hardware. It supports end-to-end reinforcement learning (RL) and offers significantly faster sampling. The library is designed to be modular, allowing easy prototyping of new training algorithms, environments, and execution plans without compromising usability or speed.

NovaSky AINov 3
Open

Step-Audio-R1

Step-Audio-R1 is an advanced audio language model developed by StepFun AI, designed to enhance audio reasoning capabilities by grounding its reasoning in acoustic features. It introduces Modality-Grounded Reasoning Distillation (MGRD), an iterative training framework that shifts the model's reasoning from textual abstractions to acoustic properties, effectively addressing the 'inverted scaling' problem where performance degrades with longer reasoning. This model has demonstrated superior performance across various audio understanding and reasoning benchmarks, surpassing models like Gemini 2.5 Pro and achieving results comparable to Gemini 3 Pro.

StepFun AINov 29
Open

SINQ

SINQ (Sinkhorn-Normalized Quantization) is a novel, fast, and high-quality quantization method designed to make any Large Language Model (LLM) smaller while preserving accuracy. It offers a plug-and-play, model-agnostic technique that delivers state-of-the-art performance for LLMs without sacrificing accuracy.

HuaweiNov 14
Open

AIBuildAI

AIBuildAI is an AI agent that autonomously constructs AI models. Given a specific task, it initiates an agent loop to analyze the problem, design models, and execute training processes, all without human intervention.

AIBuildAI
Open

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode