---
name: zhipu
description: Use when building AI applications with language models, vision models, image/video generation, or agents. Reach for this skill when implementing chat completions, function calling, streaming responses, structured output, or integrating with coding tools via the GLM Coding Plan.
metadata:
    mintlify-proj: zhipu
    version: "1.0"
---

# Z.AI Platform Skill

## Product Summary

Z.AI is a comprehensive AI platform providing access to GLM models (language, vision, multimodal), image/video generation, and agent capabilities. Agents use Z.AI to build chat applications, implement function calling, generate content, and deploy AI-powered workflows. The platform offers HTTP REST APIs, official SDKs (Python, Java), and OpenAI-compatible endpoints. Primary API endpoint: `https://api.z.ai/api/paas/v4`. Key files: API keys managed at `https://z.ai/manage-apikey/apikey-list`. SDKs: `zai-sdk` (Python), `ai.z.openapi:zai-sdk` (Java). CLI: Use curl, Python SDK, or Java SDK for API calls. Integrations: Claude Code, Cline, Roo Code, and other coding tools via MCP protocol.

## When to Use

Use this skill when:
- Building chat applications or conversational AI with GLM-5.3, GLM-5.3-Flash, or other language models
- Implementing function calling to enable models to invoke external APIs or tools
- Generating images from text prompts (GLM-Image, CogView-4)
- Creating videos from text or images (CogVideoX-3)
- Transcribing audio (GLM-ASR-2512)
- Enabling streaming responses for real-time user interaction
- Requiring structured JSON output from model responses
- Using context caching for long-context conversations
- Deploying agents with deep thinking/reasoning capabilities
- Integrating with coding tools via GLM Coding Plan
- Extracting or parsing data with models

## Quick Reference

### API Authentication
```
Authorization: Bearer YOUR_API_KEY
Base URL: https://api.z.ai/api/paas/v4
```

### Core Models
| Model | Strength | Context | Output |
|-------|----------|---------|--------|
| GLM-5.3 | Flagship, coding, reasoning | 1M | 128K |
| GLM-5.3-Flash | Multimodal, cost-effective | 1M | 128K |
| GLM-5.2 | Reasoning, versatile | 1M | 128K |
| GLM-4.7 | Agentic coding | 200K | 128K |
| GLM-4.6 | Strong coding | 200K | 128K |

### Essential Parameters
| Parameter | Type | Default | Purpose |
|-----------|------|---------|---------|
| `model` | string | — | Model ID (e.g., "glm-5.3") |
| `messages` | array | — | Conversation history with roles: "system", "user", "assistant" |
| `stream` | boolean | false | Enable streaming output |
| `temperature` | float | varies | Randomness (0.0–2.0); lower = deterministic, higher = creative |
| `max_tokens` | integer | model-dependent | Max output length in tokens |
| `thinking` | object | `{"type": "enabled"}` | Enable chain-of-thought reasoning |
| `reasoning_effort` | string | "max" | Reasoning depth: "low", "high", "max" |
| `response_format` | object | — | Set to `{"type": "json_object"}` for JSON mode |
| `tools` | array | — | Function definitions for function calling |
| `tool_choice` | string | "auto" | Control function calling strategy |

### SDK Installation
```bash
# Python
pip install zai-sdk

# Java (Maven)
<dependency>
    <groupId>ai.z.openapi</groupId>
    <artifactId>zai-sdk</artifactId>
    <version>0.3.5</version>
</dependency>
```

### Basic Chat Call (Python)
```python
from zai import ZaiClient

client = ZaiClient(api_key="YOUR_API_KEY")
response = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)
```

## Decision Guidance

### When to Use X vs Y

| Scenario | Use | Reason |
|----------|-----|--------|
| **Streaming vs Non-Streaming** | Streaming (`stream=True`) | Real-time UX, long responses, chat apps |
| | Non-Streaming (`stream=False`) | Simple scripts, batch processing, when full response needed at once |
| **Temperature Control** | Low (0.2–0.5) | Factual Q&A, code generation, deterministic output |
| | High (0.7–1.0) | Creative writing, brainstorming, diverse responses |
| **GLM-5.3 vs GLM-5.3-Flash** | GLM-5.3 | Complex reasoning, deep coding tasks, max quality |
| | GLM-5.3-Flash | Multimodal input, cost-sensitive, faster inference |
| **Function Calling vs Direct API** | Function Calling | Model decides when to call tools, multi-step workflows |
| | Direct API | Simple single-step operations, no tool invocation needed |
| **JSON Mode vs Free Text** | JSON Mode (`response_format`) | Structured data extraction, API responses, validation |
| | Free Text | Natural language, creative content, flexible format |
| **Thinking Enabled vs Disabled** | Enabled (`thinking.type: "enabled"`) | Complex reasoning, debugging, multi-step problems |
| | Disabled (GLM-5.2 only) | Simple tasks, speed priority, cost optimization |
| **Context Caching** | Enable | Long conversations, repeated context, cost savings |
| | Disable | Short single-turn interactions, no repeated content |

## Workflow

### 1. Set Up Authentication
- Create API key at `https://z.ai/manage-apikey/apikey-list`
- Store as environment variable: `export ZAI_API_KEY=your-key`
- Or pass directly: `ZaiClient(api_key="your-key")`

### 2. Choose Model and Capabilities
- Select model based on task (GLM-5.3 for complex reasoning, GLM-5.3-Flash for multimodal)
- Decide on features: streaming, function calling, structured output, thinking
- Check context window and max output limits for your use case

### 3. Construct Messages
- Start with optional system message to set tone/role
- Add user messages and assistant responses for multi-turn conversations
- Preserve message history for context in long conversations

### 4. Configure Parameters
- Set `temperature` based on creativity needs (low for facts, high for creativity)
- Enable `stream=True` for real-time responses
- Add `tools` array if using function calling
- Set `response_format={"type": "json_object"}` for structured output
- Configure `thinking` and `reasoning_effort` for complex tasks

### 5. Make API Call
- Use SDK (Python/Java) or HTTP API (curl/REST)
- Handle streaming chunks if `stream=True`
- Parse response: `response.choices[0].message.content`

### 6. Handle Function Calls (if applicable)
- Check `response.choices[0].message.tool_calls`
- Execute each function locally
- Return results in `tool` role messages
- Make follow-up call to get final answer

### 7. Verify and Log
- Check `finish_reason` (should be "stop" or "tool_calls")
- Log token usage from `response.usage`
- Validate JSON output if using `response_format`
- Monitor rate limits and quota consumption

## Common Gotchas

- **GLM-5.3 requires thinking enabled**: Cannot disable reasoning; set `reasoning_effort` to "low" for lightweight thinking instead
- **API key in code**: Never hardcode API keys; use environment variables or secure vaults
- **Streaming chunks without content**: Check `if chunk.choices[0].delta.content` before accessing; not all chunks contain content
- **Token counting**: Tokens ≠ words; 1 token ≈ 0.75 English words or 1.5 Chinese characters; use tokenizer API to estimate
- **Context window overflow**: Input + output must fit within model's context window (e.g., 1M for GLM-5.3); truncate history if needed
- **Function calling requires tool_calls handling**: Don't forget to return tool results in messages; model won't proceed without them
- **JSON mode strictness**: Model must return valid JSON; provide clear schema in system message to avoid parse errors
- **Rate limits and quotas**: Check `https://z.ai/manage-apikey/rate-limits`; 5-hour and weekly limits apply; peak hours (Mon–Fri 14:00–18:00 SGT) consume 3× quota for GLM-5.3
- **Streaming with function calls**: Use `tool_stream=true` to get real-time tool parameters during streaming
- **Preserved thinking**: When using `clear_thinking: false`, return reasoning_content blocks exactly as generated; reordering breaks cache
- **Model deprecation**: Always check latest model availability; older models may be taken offline
- **OpenAI SDK compatibility**: Base URL differs: `https://api.z.ai/api/paas/v4/` for OpenAI SDK vs standard endpoint
- **GLM Coding Plan endpoint**: Requires separate configuration; use dedicated endpoint from plan dashboard, not standard API endpoint

## Verification Checklist

Before submitting work:
- [ ] API key is valid and has sufficient balance
- [ ] Model name is correct and supported (check docs for latest)
- [ ] Messages array is properly formatted with role/content pairs
- [ ] If using function calling, all tool definitions include name, description, and parameters
- [ ] If using function calling, tool results are returned in messages with correct tool_call_id
- [ ] If using streaming, code handles chunks without content gracefully
- [ ] If using JSON mode, response_format is set and schema is clear in system message
- [ ] If using thinking, reasoning_effort is set appropriately (not just enabled)
- [ ] Temperature and top_p are not both set (use one or the other)
- [ ] max_tokens is reasonable for the task (not too low to truncate, not too high to waste tokens)
- [ ] Error handling covers API errors (401 auth, 429 rate limit, 1200+ business codes)
- [ ] Token usage is logged for cost tracking
- [ ] Context window is not exceeded (check model's max context)
- [ ] For production: API key is in environment variable, not hardcoded
- [ ] For streaming: finish_reason is checked to detect completion
- [ ] For long conversations: message history is pruned or cached to avoid overflow

## Resources

- **Comprehensive navigation**: https://docs.z.ai/llms.txt
- **API Reference**: https://docs.z.ai/api-reference
- **Python SDK Guide**: https://docs.z.ai/guides/develop/python/introduction
- **Quick Start**: https://docs.z.ai/guides/overview/quick-start
- **Core Parameters**: https://docs.z.ai/guides/overview/concept-param
- **Function Calling**: https://docs.z.ai/guides/capabilities/function-calling
- **Streaming**: https://docs.z.ai/guides/capabilities/streaming
- **Structured Output**: https://docs.z.ai/guides/capabilities/struct-output
- **GLM-5.3 Model**: https://docs.z.ai/guides/llm/glm-5.3
- **Error Codes**: https://docs.z.ai/api-reference/api-code
- **Rate Limits**: https://docs.z.ai/api-reference/rate-limit

---

> For additional documentation and navigation, see: https://docs.z.ai/llms.txt