Tencent

2026Tencent

Youtu-VL-4B-Instruct-GGUF v1.0

Open Source
CV

Youtu-VL-4B-Instruct-GGUF is a lightweight yet robust Vision-Language Model (VLM) developed by Tencent. Built upon the Youtu-LLM with 4 billion parameters, it introduces the Vision-Language Unified Autoregressive Supervision (VLUAS) paradigm, enhancing visual perception and multimodal understanding. This model excels in both vision-centric and general multimodal tasks without the need for task-specific modules.

Comprehensive Vision-Centric Capabilities
Promising Performance with High Efficiency
Unified Autoregressive Supervision
PricingFree
Version1.0
2026Tencent

CL-bench v1.0

Open Source
Coding
LLM

CL-bench is a benchmark designed to evaluate language models' ability to learn from complex, context-dependent tasks. It comprises 500 complex contexts, 1,899 tasks, and 31,607 verification rubrics, all crafted by experienced domain experts.

500 complex contexts
1,899 tasks
31,607 verification rubrics
PricingFree
Version1.0
2026Tencent

TencentDB Agent Memory vLatest (May 2026)

Paid
Featured
Agent Memory
Long-term Memory

TencentDB Agent Memory is a fully local, open-source long-term memory system for AI agents built by Tencent. Using a 4-tier progressive pipeline with symbolic short-term memory (Mermaid canvas) and layered long-term memory (L0→L3 persona hierarchy), it cuts token usage by 61% and improves agent task success by 51% when integrated with OpenClaw — with zero external API dependencies.

Symbolic short-term memory — offloads verbose tool logs to external files and condenses in-context state into compact Mermaid symbol graphs with node_id tracing; cuts WideSearch token usage by 61.38% and improves pass rate from 33% to 50% (relative +51.52%)
4-tier long-term memory pipeline — L0 raw conversation → L1 atomic facts → L2 scenario blocks → L3 user persona; layered storage with full drill-down traceability (no irreversible compression); raises PersonaMem accuracy from 48% to 76%
Fully local, zero external API dependencies — dual-layer storage (database for facts/logs + Markdown for personas/scenes); integrates as an OpenClaw plugin and Hermes skillpack; SWE-bench improvement +9.93% with 33% token savings
PricingFree (Open Source)
VersionLatest (May 2026)
2026Tencent

WorkBuddy Bench v1.0

Open Source
code
agent

WorkBuddy Bench is a multi-domain evaluation suite designed to test coding agents on realistic, complex tasks across software engineering, web development, office workflows, and security. It features 260 contamination-resistant tasks reverse-engineered from real-world commits and business scenarios to ensure agents are evaluated on genuine problem-solving capabilities rather than memorized data.

260 contamination-resistant tasks across Code, Web, Office, and Security domains
Tasks reverse-engineered from real-world commits, PRs, and business scenarios
Unified task-directory format with sandboxed Dockerized environments
PricingFree
Version1.0
2026Tencent

AngelSpec v0.1.0

Open Source
Speculative Decoding
LLM Inference

AngelSpec is a unified, torch-native training framework designed for speculative decoding, supporting both autoregressive Multi-Token Prediction (MTP) and block-parallel drafting architectures. It enables independent scaling of inference and training through a disaggregated architecture, facilitating high-performance model deployment.

Unified training pipeline for 6 draft architectures (DFly, DFlash, DFlare, Eagle3, DSpark, MTP)
Disaggregated training architecture with independent scaling for inference and optimization
Support for long-context training up to 128k tokens via Ulysses sequence parallelism
PricingFree
Version0.1.0
2026Tencent

HY-1.8B-2Bit v1.0

Open Source
LLMs
ML

HY-1.8B-2Bit is a 2-bit quantized large language model developed by Tencent's AngelSlim project. It utilizes Quantization-Aware Training (QAT) on the Hunyuan-1.8B-Instruct backbone, achieving high performance with significantly reduced model size.

2-bit quantization for efficient deployment
Maintains 96% of full-precision model performance
Dual Chain-of-Thought (Dual-CoT) strategy for flexible reasoning
Version1.0
RegionChina
2026Tencent

High Performance LLM Inference Operator Library

Open Source
Infrastructure

The High Performance LLM Inference Operator Library is an open-source project developed by Tencent, designed to optimize the inference performance of large language models (LLMs). It provides a set of operators and tools that enhance the efficiency and scalability of LLM deployments, making it easier for developers to integrate and utilize LLMs in various applications.

Optimized inference operators for LLMs
Scalable deployment support
Open-source and community-driven
PricingFree
RegionChina
2026Tencent

Cube Sandbox by Tencent v0.1.0

Paid
sandbox
ai-agents

Cube Sandbox is an open-source high-performance sandbox runtime for AI agents built on RustVMM and KVM. It achieves sub-60ms cold starts, under 5MB memory per instance, hardware-level kernel isolation, and supports thousands of concurrent sandboxes per node. 100% E2B SDK compatible with zero code changes to migrate.

Sub-60ms cold start — 2.5-50x faster than alternatives
Under 5MB memory per instance with hardware-level kernel isolation
100% E2B SDK compatible — zero code changes needed to migrate
PricingOpen Source
Version0.1.0
2025Tencent

HY-MT1.5 v1.5

Open Source
ML
CV

HY-MT1.5 is a multilingual machine translation model developed by Tencent, available in 1.8B and 7B parameter versions. It supports translation across 33 languages and 5 ethnic and dialect variations, optimized for both on-device and cloud deployment.

Supports 33 languages and 5 ethnic and dialect variations
Optimized for on-device and cloud deployment
Includes terminology intervention, contextual translation, and formatted translation capabilities
PricingFree
Version1.5
2026Tencent

Hy3 v3

Open Source
Large Language Model
Mixture-of-Experts

Hy3 is a powerful 295B-parameter Mixture-of-Experts (MoE) large language model developed by Tencent. It features 21B active parameters per token, a 256K context window, and is designed to deliver high-performance reasoning, agentic workflows, and long-context processing at lower compute costs.

295B total parameters with 21B active parameters per token
256K context window for long-form document comprehension
Advanced agentic capabilities with support for tool-call orchestration
PricingFree
Version3

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode