Models
GLM-5.3-Flash
- Native visual capabilities enable the model to observe interfaces, rendering results, and interaction feedback—creating a closed loop across code, browsers, and GUIs.
- Efficient hybrid architecture: Combines linear and sparse attention with 320B total parameters and 18B activated, significantly reducing compute and KV-cache requirements.
- Beyond coding: Supports office document and financial research workflows, autonomously breaking down goals, using tools, and refining outputs. Learn more in our documentation.*
GLM-5.3
- Stronger Coding Capabilities: GLM-5.3 delivers a significant improvement in coding capabilities, achieving a 50% gain over GLM-5.2 on Z.ai Code Bench and reaching state-of-the-art (SOTA) performance among open-source models on public benchmarks, including Terminal Bench 3.0.
- Emergent Cybersecurity Capabilities: GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery. In collaboration with multiple cybersecurity teams, it has been tested on real-world targets and has identified a total of 2,436 vulnerabilities, including 1,097 medium- and high-severity vulnerabilities.Learn more in our documentation.*
GLM-5.2
- Supports 1M lossless context, significantly improving long-horizon task capabilities and reducing context drift and goal forgetting in complex tasks
- Achieves open-source SOTA performance on coding and long-horizon task benchmarks, delivering more stable results in complex system engineering and deep debugging
- Significantly improves the real-world developer experience, with more reliable project-level context handling, adherence to engineering standards, and multi-platform development. Learn more in our documentation.*
GLM-5.1
- Designed for long-horizon tasks, GLM-5.1 can work independently for up to 8 hours in a single run, enabling a full loop from planning and execution to iterative refinement and final delivery.
- It demonstrates stronger engineering intelligence across autonomous planning, sustained execution, bug fixing, and strategy iteration, while achieving comprehensive capability alignment with Claude Opus 4.6. Built with multi-turn SFT, RL, and a process-quality evaluation framework, GLM-5.1 further improves stability, consistency, and tool use over extended tasks. Learn more in our documentation.*
GLM-5
- Designed for complex system engineering and long-range Agent tasks, GLM-5 shifts the paradigm from coding to engineering, demonstrating strong deep-reasoning performance in backend architecture, complex algorithms, and stubborn bug fixing.
- It directly benchmarks against Claude Opus 4.5 in code-logic density and systems-engineering capability, and integrates DeepSeek Sparse Attention for higher token efficiency while preserving long-context quality.Learn more in our documentation.*
GLM-OCR
- We’ve launched GLM-OCR, a compact and high-performance optical character recognition model powered by the self-developed CogViT and GLM-0.5B encoder-decoder architecture, enabling efficient cross-modal alignment through its dedicated connection layer.
- The update leverages CLIP pre-training on billions of image-text pairs to deliver robust visual semantic understanding and key token extraction capabilities, while maintaining a lightweight design for fast inference. Learn more in our documentation.*
GLM-4.7-Flash
- We’ve launched GLM-4.7-Flash, a lightweight and efficient model designed as the free-tier version of GLM-4.7, delivering strong performance across coding, reasoning, and generative tasks with low latency and high throughput.
- The update brings competitive coding capabilities at its scale, offering best-in-class general abilities in writing, translation, long-form content, role play, and aesthetic outputs for high-frequency and real-time use cases. Learn more in our documentation.*
GLM-Image
- We’ve launched GLM-Image, a state-of-the-art image generation model built on a multimodal architecture and fully trained on domestic chips, combining autoregressive semantic understanding with diffusion-based decoding to deliver high-quality, controllable visual generation.
- The update significantly enhances performance in knowledge-intensive scenarios, with more stable and accurate text rendering inside images, making GLM-Image especially well suited for commercial design, educational illustrations, and content-rich visual applications.Learn more in our documentation.*
GLM-4.7
- We’ve released GLM-4.7, our foundation model with significant improvements in coding, reasoning, and agentic capabilities. It delivers more reliable code generation, stronger long-context understanding, and improved end-to-end task execution across real-world development workflows.
- The update brings open-source SOTA performance on major coding and reasoning benchmarks, enhanced agentic coding for goal-driven, multi-step tasks, and improved front-end and document generation quality. Learn more in our documentation.*
AutoGLM-Phone-Multilingual
- We’ve launched AutoGLM-Phone-Multilingual, our latest multimodal mobile automation framework that understands screen content and executes real actions through ADB. It enables natural-language task execution across 50+ mainstream apps, delivering true end-to-end mobile control.
- The update introduces multilingual support (English & Chinese), enhanced workflow planning capabilities, and improved task execution reliability. Learn more in our documentation.*
GLM-ASR-2512
- We’ve launched GLM-ASR-2512, our ASR model, delivering industry-leading accuracy with a Character Error Rate of just 0.0717, and significantly improved performance across real-world multilingual and accent-rich scenarios.
- The update introduces enhanced custom dictionary support and expanded specialized terminology recognition. Learn more in our documentation.*
GLM-4.6V
- We’re excited to introduce GLM-4.6V, Z.ai’s latest iteration in multimodal large language models. This version enhances vision understanding, achieving state-of-the-art performance in tasks involving images and text.
- The update also expands the context window to 128K, enabling more efficient processing of long inputs and complex multimodal tasks. Learn more in our documentation.*
GLM-4.6
- We’ve launched GLM-4.6, the flagship coding model, showcasing enhanced performance in both public benchmarks and real-world programming tasks, making it the leading coding model in China.
- The update also expands the context window to 200K, improving its ability to handle longer code and complex agent tasks. Learn more in our documentation.*
GLM-4.5V
- We’ve launched GLM-4.5V, a 100B-scale open-source vision reasoning model, supporting a broad range of visual tasks including video understanding, visual grounding, GUI agents and etc.
- The update also adds a new thinking mode. Learn more in our documentation.*
GLM Slide/Poster Agent(beta)
- We’ve launched GLM Slide/Poster Agent, an AI-powered creation agent that combines information retrieval, content structuring, and visual layout design to generate professional-grade slides and posters from natural language instructions.
- The update also brings a seamless integration of content generation with design conventions. Learn more in our documentation.*
GLM-4.5 Series
- We’ve launched GLM-4.5, our latest native agentic LLM, delivering doubled parameter efficiency and strong reasoning, coding, and agentic capabilities.
- It also offers seamless one-click compatibility with the Claude Code framework. Learn more in our documentation.*
CogVideoX-3
- We’ve launched CogVideoX-3, an incremental upgrade to our video generation model with improved quality and new features.
- It adds support for start and end frame synthesis. Learn more in our documentation.*