Zhipu AI
GLM-5 v5
GLM-5 is Z.AI's latest flagship foundation model, designed for complex system engineering and long-range agentic tasks. It features a parameter scale of 744 billion, an expanded pre-training dataset of 28.5 trillion tokens, and integrates DeepSeek Sparse Attention for efficient long-context processing.
GLM-Image v1.0
GLM-Image is Z.AI's flagship image generation model that combines an autoregressive module with a diffusion decoder. This hybrid architecture excels in generating high-quality, knowledge-intensive images, such as posters, presentations, and educational diagrams.
GLM-4.7 v4.7
GLM-4.7 is Z.AI's latest flagship foundation model, offering significant improvements in coding, reasoning, and agentic capabilities. It delivers more reliable code generation, stronger long-context understanding, and enhanced end-to-end task execution across real-world development workflows.
GLM-OCR
GLM-OCR is a lightweight professional OCR model with parameters as small as 0.9B, yet it achieves state-of-the-art performance across multiple capabilities. It sets a new benchmark for document parsing with its "small size and high accuracy." Key features include: - Performance SOTA: Scored 94.62 points to top OmniDocBench V1.5 and achieved current best performance across multiple mainstream document understanding benchmarks including tables and formulas at launch. - Optimized for Real-World Scenarios: Delivers stable, leading accuracy in complex environments like code documentation, intricate tables, and stamp recognition. Maintains exceptional recognition precision even with complex layouts, diverse fonts, or mixed text-image content. - Efficient and Cost-Effective: With just 0.9B parameters, supports VLLM and SGLang deployment, significantly reducing inference latency and computational overhead.
ZCode v3.2.2
ZCode is a dedicated Agentic Development Environment (ADE) optimized for the GLM-5.2 model, designed to handle complex, long-horizon software engineering tasks. It provides a centralized interface for planning, coding, reviewing, and deploying projects, featuring native tools for terminal integration, Git version control, and multi-agent collaboration.
GLM-TTS v1.0
GLM-TTS is a high-quality text-to-speech (TTS) synthesis system based on large language models, supporting zero-shot voice cloning and streaming inference. It utilizes a two-stage architecture combining a language model for speech token generation and a Flow Matching model for waveform synthesis. By introducing a Multi-Reward Reinforcement Learning framework, GLM-TTS significantly improves the expressiveness of generated speech, achieving more natural emotional control compared to traditional TTS systems.
GLM-4.6V v4.6V
GLM-4.6V is a 106-billion parameter vision language model designed to process images, videos, and tools as primary inputs for agents. It extends the training context window to 128,000 tokens, enabling the processing of approximately 150 pages of dense documents, 200 slide pages, or one hour of video in a single pass. The model introduces native multimodal function calling, allowing direct processing of images, screenshots, and document pages as tool parameters, thereby bridging the gap between visual perception and executable action for multimodal agents.