Z.ai · Model
GLM 5 Turbo pricing, context window and specs
GLM 5 Turbo costs $1.20 per million input tokens and $4 per million output tokens, and reads up to 202,752 tokens per request. By blended price it is number 4 of Z.ai’s 7 current models here, counting from the cheapest.
- Input / 1M
- $1.20
- Output / 1M
- $4
- Cached / 1M
- $0.24
- Context
- 202,752
- Max output
- 131,072
- Released
- 2026-03-15
Estimate your cost with GLM 5 TurboCompare all Z.ai modelsCount tokens for GLM 5 Turbo
Pricing
GLM 5 Turbo API pricing
| Price per 1M tokens | Input | Output |
|---|---|---|
| Standard | $1.20 | $4 |
| Cached input (read) | $0.24 | – |
Output costs 3.3× the input price, so long replies drive the bill more than long prompts. If your requests share a long fixed prefix, such as a system prompt or tool definitions, caching cuts that part to 20% of the normal input price.
Examples
What GLM 5 Turbo costs in practice
| Workload | Tokens in / out | Requests/day | Per request | Per month |
|---|---|---|---|---|
| Support chatbot50% cached | 1,500 / 400 | 1,000 | $0.00268 | $81.52 |
| RAG app20% cached | 6,000 / 500 | 500 | $0.00805 | $122.40 |
| Coding agent80% cached | 40,000 / 2,000 | 200 | $0.0253 | $153.79 |
| Document summariser | 8,000 / 600 | 300 | $0.012 | $109.50 |
Change any number in the cost calculator.
Context
Context window
GLM 5 Turbo accepts up to 202,752 tokens per request, roughly 152,064 English words or 304 pages. That budget is shared between your prompt and the reply, and a single reply is capped at 131,072 tokens.
Filling the whole window with one prompt costs about $0.24 at list price, so for repeated long-document work, retrieval (sending only the relevant parts) or prompt caching is usually much cheaper.
Features
Capabilities
- Image input
- No
- Tool / function calling
- Yes
- Reasoning mode
- Yes
- Prompt caching
- Yes
- Structured output
- Yes
- Open weights
- No
- Accepts
- text
- Knowledge cutoff
- Not published
As declared by the provider’s API. It shows a feature exists, not how well it works.
API
GLM 5 Turbo model ID
- OpenRouter
z-ai/glm-5-turbo
Use the OpenRouter ID when calling it through OpenRouter; Z.ai’s own API may name it differently. See the model ID reference for the providers we document.
Z.ai
Where GLM 5 Turbo sits in Z.ai’s line-up
Ranked by blended price (three parts input to one part output, $1.90 per 1M for GLM 5 Turbo), it is number 4 of Z.ai’s 7 current models here, counting from the cheapest. One step down, GLM 5.3 FlashX costs $0.59 blended (69% less). Unlike GLM 5 Turbo, it accepts images and video and reads up to 1,048,576 tokens (202,752 here). One step up, GLM 5V Turbo costs $1.90, the same. Unlike GLM 5 Turbo, it accepts images and video. See every Z.ai model and price
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| GLM 5.3 Flash | $0.15 | $0.50 | 1,048,575 | 2026-08-26 |
| GLM 4.6V | $0.30 | $0.90 | 131,072 | 2025-12-08 |
| GLM 5.3 FlashX | $0.37 | $1.25 | 1,048,576 | 2026-09-18 |
| GLM 5 Turbo (this page) | $1.20 | $4 | 202,752 | 2026-03-15 |
| GLM 5V Turbo | $1.20 | $4 | 202,752 | 2026-04-01 |
| GLM 5.3 | $1.40 | $4.40 | 1,048,576 | 2026-08-18 |
| GLM 5.3 Prime | $2.80 | $8.80 | 1,000,000 | 2026-09-23 |
Z.ai’s current models with a page here, cheapest first. Older and superseded models are listed on the provider page and in the comparison table.
Save
Cheaper alternatives
- GPT-6 LunaOpenAI89% cheaper
- Claude Haiku 5.5Anthropic89% cheaper
- Qwen3.8 Omni FlashQwen88% cheaper
- DeepSeek V4 Pro 0423DeepSeek86% cheaper
- GLM 4.6VZ.ai76% cheaper
Current models from major providers, one per provider, that keep tool calling, a reasoning mode, a usable context window and output limit. Compared by blended price.
Compare
Similarly priced models
- DeepSeek V4 Pro 0813DeepSeek$1.32 / $3.96
- Muse Spark 1.3Meta$1.25 / $4.25
- GPT-5.4 MiniOpenAI$0.75 / $4.50
- Gemini 3.8 FlashGoogle$0.75 / $3.75
- Kimi K2.7 CodeMoonshot AI$0.6712 / $3.35
Current models from other major providers, closest blended price.
FAQ
GLM 5 Turbo questions
How much does GLM 5 Turbo cost?
GLM 5 Turbo costs $1.20 per million input tokens and $4 per million output tokens, with cached input at $0.24. A typical chatbot reply (1,500 tokens in, 400 out) costs about $0.0034. Use the LLM cost calculator for your own workload.
What is GLM 5 Turbo’s context window?
202,752 tokens, about 152,064 English words or 304 printed pages, shared between your prompt and the reply. A single reply can be up to 131,072 tokens.
Does GLM 5 Turbo support prompt caching and batch requests?
Yes, cached input is billed at $0.24 per million tokens (20% of the normal input price). No batch price is published for it.
What are cheaper alternatives to GLM 5 Turbo?
With the same essentials (tool calling), the cheapest options are GPT-6 Luna (89% cheaper), Claude Haiku 5.5 (89% cheaper), Qwen3.8 Omni Flash (88% cheaper). Cheaper doesn’t mean equivalent, so test them on your own prompts.
What is the API model ID for GLM 5 Turbo?
On OpenRouter it is z-ai/glm-5-turbo. Z.ai’s own API may use a different name, so check its model list before you deploy. Pass the ID as the model value in each request.
Is GLM 5 Turbo open source?
No, GLM 5 Turbo is only available through Z.ai’s API and partners; its weights aren’t published.