Google · Model
Gemma 4 31B pricing, context window and specs
Gemma 4 31B costs $0.09 per million input tokens and $0.34 per million output tokens, and reads up to 262,144 tokens per request. By blended price it is number 2 of Google’s 5 current models here, counting from the cheapest.
- Input / 1M
- $0.09
- Output / 1M
- $0.34
- Cached / 1M
- $0.05
- Context
- 262,144
- Max output
- 16,384
- Released
- 2026-04-02
Estimate your cost with Gemma 4 31BCompare all Google modelsCount tokens for Gemma 4 31B
Pricing
Gemma 4 31B API pricing
| Price per 1M tokens | Input | Output |
|---|---|---|
| Standard | $0.09 | $0.34 |
| Cached input (read) | $0.05 | – |
Output costs 3.8× the input price, so long replies drive the bill more than long prompts. If your requests share a long fixed prefix, such as a system prompt or tool definitions, caching cuts that part to 56% of the normal input price.
Examples
What Gemma 4 31B costs in practice
| Workload | Tokens in / out | Requests/day | Per request | Per month |
|---|---|---|---|---|
| Support chatbot50% cached | 1,500 / 400 | 1,000 | $0.000241 | $7.33 |
| RAG app20% cached | 6,000 / 500 | 500 | $0.000662 | $10.07 |
| Coding agent80% cached | 40,000 / 2,000 | 200 | $0.003 | $18.25 |
| Document summariser | 8,000 / 600 | 300 | $0.000924 | $8.43 |
Change any number in the cost calculator.
Context
Context window
Gemma 4 31B accepts up to 262,144 tokens per request, roughly 196,608 English words or 393 pages. That budget is shared between your prompt and the reply, and a single reply is capped at 16,384 tokens.
Filling the whole window with one prompt costs about $0.0236 at list price, so for repeated long-document work, retrieval (sending only the relevant parts) or prompt caching is usually much cheaper.
Features
Capabilities
- Image input
- Yes
- Tool / function calling
- Yes
- Reasoning mode
- Yes
- Prompt caching
- Yes
- Structured output
- Yes
- Open weights
- Yes
- Accepts
- image, text, video
- Knowledge cutoff
- Not published
As declared by the provider’s API. It shows a feature exists, not how well it works.
API
Gemma 4 31B model ID
- OpenRouter
google/gemma-4-31b-it- Open weights
- google/gemma-4-31B-it
Use the OpenRouter ID when calling it through OpenRouter; Google’s own API may name it differently. See the model ID reference for the providers we document. See how much VRAM Gemma 4 31B needs to run locally.
Where Gemma 4 31B sits in Google’s line-up
Ranked by blended price (three parts input to one part output, $0.1525 per 1M for Gemma 4 31B), it is number 2 of Google’s 5 current models here, counting from the cheapest. One step down, Gemma 4 26B A4B costs $0.1069 blended (30% less). One step up, Gemini 3.5 Flash Lite costs $0.85 (5.6× as much). Unlike Gemma 4 31B, it accepts files and audio, reads up to 1,048,576 tokens (262,144 here) and caps a reply at 65,536 tokens (16,384 here). See every Google model and price
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| Gemma 4 26B A4B | $0.0675 | $0.225 | 262,144 | 2026-04-03 |
| Gemma 4 31B (this page) | $0.09 | $0.34 | 262,144 | 2026-04-02 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | 1,048,576 | 2026-07-21 |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1,048,576 | 2026-09-02 |
| Gemini 3.1 Pro Preview | $2 | $12 | 1,048,576 | 2026-02-19 |
Google’s current models with a page here, cheapest first. Older and superseded models are listed on the provider page and in the comparison table.
Save
Cheaper alternatives
- Qwen3.5-9BQwen26% cheaper
- Seed 1.6 FlashByteDance Seed14% cheaper
Current models from major providers, one per provider, that keep image input, tool calling, a reasoning mode, a usable context window and output limit. Compared by blended price.
Compare
Similarly priced models
- Ministral 3 8B 2512Mistral$0.15 / $0.15
- Nemotron 3 SuperNVIDIA$0.08 / $0.45
- Seed-2.0-MiniByteDance Seed$0.10 / $0.40
- GPT-6 LunaOpenAI$0.10 / $0.50
- Claude Haiku 5.5Anthropic$0.10 / $0.50
Current models from other major providers, closest blended price.
FAQ
Gemma 4 31B questions
How much does Gemma 4 31B cost?
Gemma 4 31B costs $0.09 per million input tokens and $0.34 per million output tokens, with cached input at $0.05. A typical chatbot reply (1,500 tokens in, 400 out) costs about $0.000271. Use the LLM cost calculator for your own workload.
What is Gemma 4 31B’s context window?
262,144 tokens, about 196,608 English words or 393 printed pages, shared between your prompt and the reply. A single reply can be up to 16,384 tokens.
Does Gemma 4 31B support prompt caching and batch requests?
Yes, cached input is billed at $0.05 per million tokens (56% of the normal input price). No batch price is published for it.
What are cheaper alternatives to Gemma 4 31B?
With the same essentials (image input and tool calling), the cheapest options are Qwen3.5-9B (26% cheaper), Seed 1.6 Flash (14% cheaper). Cheaper doesn’t mean equivalent, so test them on your own prompts.
What is the API model ID for Gemma 4 31B?
On OpenRouter it is google/gemma-4-31b-it. Google’s own API may use a different name, so check its model list before you deploy. Pass the ID as the model value in each request.
Is Gemma 4 31B open source?
Yes, Gemma 4 31B has open weights, so you can run it yourself or choose between several hosting providers. The prices here are a typical hosted price.