Skip to main content
Z.AI offers a variety of models and agents to meet the needs of different scenarios. Choosing the right model can help you complete tasks more efficiently.

GLM-5.3

Open-source SOTA coding capabilities, emerging cybersecurity capabilities

GLM-5V-Turbo

Multimodal coding model, specializing in visual programming

GLM-Image

Supports text-to-image generation, achieving open-source state-of-the-art in complex scenarios

Models, Agents and Tools

To help you find the best fit for your use case, we’ve created a table outlining the core features and strengths of each model in the Z.AI family.
If you need to get pricing information, please go directly to Pricing.

Text Models

Our model matrix includes text models with built-in reasoning capabilities, as well as vision-language models (VLMs) that extend the same reasoning power to multimodal understanding.

Vision Models

Visual models process images or videos for recognition and analysis.

Built-in Tools

A suite of built-in tools designed to streamline workflows and boost productivity.

Image Generation Models

Image Generation Models learn from massive image data to automatically generate high-quality images from text.

Video Generation Models

Video Generation Models turn text, images, or clips into dynamic video content, accelerating creativity for film, virtual avatars, animation, and marketing.

Audio Models

Audio models are a class of multimodal models that process audio and video signals, enabling the understanding, generation, or editing of audiovisual content.

Agents

A set of ready-made agents empower users to create and communicate effortlessly.