Z.AI offers a variety of models and agents to meet the needs of different scenarios. Choosing the right model can help you complete tasks more efficiently.
Our model matrix includes text models with built-in reasoning capabilities, as well as vision-language models (VLMs) that extend the same reasoning power to multimodal understanding.
Model
Strength
Language
Context
Resource
GLM-5.3
Claude Fable 5-level coding and agent capabilities Stronger in long-horizon, complex tasks
Video Generation Models turn text, images, or clips into dynamic video content, accelerating creativity for film, virtual avatars, animation, and marketing.
Model
Strength
Language
Resolution
Resource
CogVideoX-3
Significant improvements in image quality, stability, and physical realism simulation
Audio models are a class of multimodal models that process audio and video signals, enabling the understanding, generation, or editing of audiovisual content.
Model
Strength
Multimodal Support
Resource
GLM-ASR-2512
- CER as low as 0.0717 - Support user-defined vocabularies - Support multiple mainstream languages and dialects