Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Provider

Z.ai API pricing: every model compared

Z.ai offers 16 models through its API in our data, priced from $0.0605 to $2.80 per million input tokens. The newest is GLM 5.3 Prime, first listed on 2026-09-23.

Models
16
Cheapest input
$0.0605
Priciest input
$2.80
Largest context
1,048,576
With caching
15 of 16
Open weights
12 of 16

Filter and sort Z.ai models in the comparison table

Cheapest

GLM 4.7 FlashZ.ai$0.0605 in · $0.40 out

Largest context

GLM 5.2Z.ai1,048,576 tokens

Newest

GLM 5.3 PrimeZ.ai2026-09-23

Line-up

Every Z.ai model and price

ModelInput / 1MOutput / 1MContextReleased
GLM 5.3 Prime$2.80$8.801,000,0002026-09-23
GLM 5.3 FlashX$0.37$1.251,048,5762026-09-18
GLM 5.3 Flash$0.15$0.501,048,5752026-08-26
GLM 5.3$1.40$4.401,048,5762026-08-18
GLM 5.2$1.40$4.401,048,5762026-06-16
GLM 5.1$0.966$3.036200,0002026-04-07
GLM 5V Turbo$1.20$4202,7522026-04-01
GLM 5 Turbo$1.20$4202,7522026-03-15
GLM 5$0.60$1.92198,0002026-02-11
GLM 4.7 Flash$0.0605$0.40131,0722026-01-19
GLM 4.7$0.60$2.20202,7522025-12-22
GLM 4.6V$0.30$0.90131,0722025-12-08
GLM 4.6$0.43$1.75198,0002025-09-30
GLM 4.5V$0.60$1.8065,5362025-08-11
GLM 4.5$0.60$2.20131,0722025-07-25
GLM 4.5 Air$0.13$0.85131,0722025-07-25

Prices in US dollars per 1M tokens, newest first. Release dates are when our source first listed the model.

Overview

Z.ai’s pricing at a glance

The spread between Z.ai’s cheapest and most expensive model is large: GLM 5.3 Prime costs 46× more per input token than GLM 4.7 Flash. Picking the smallest model that does the job well is usually the biggest saving available, ahead of any discount.

15 of 16 of the line-up has a cached-input price, which helps chatbots and agents that resend the same instructions. 2 of 16 can run through a batch API at a discount, for work that can wait up to a day. 5 of 16 accept images as input.

FAQ

Z.ai API questions

How much does the Z.ai API cost?

Z.ai's 16 models range from $0.0605 to $2.80 per million input tokens, and output costs more than input on almost every model. The exact bill depends on your token volumes, which you can estimate in the LLM cost calculator.

What is Z.ai's cheapest model?

By blended price (three parts input to one part output) it is GLM 4.7 Flash, at $0.0605 input and $0.40 output per million tokens. The most expensive is GLM 5.3 Prime.

Which Z.ai model has the largest context window?

GLM 5.2, with 1,048,576 tokens per request.

Does Z.ai offer prompt caching or batch discounts?

In our data, 15 of 16 of Z.ai's models have a published cached-input price and 2 of 16 have a published batch price. Both can cut costs substantially for repeated prompts or work that can wait.

Updated

Sources: OpenRouter models API and LiteLLM model prices and context windows, checked daily. Not affiliated with Z.ai. Confirm critical numbers on the provider’s pricing page. See our methodology.