OpenAI · Model
GPT-4o-mini pricing, context window and specs
GPT-4o-mini costs $0.15 per million input tokens and $0.60 per million output tokens, and reads up to 128,000 tokens per request. By blended price it is number 3 of OpenAI’s 10 current models here, counting from the cheapest.
- Input / 1M
- $0.15
- Output / 1M
- $0.60
- Cached / 1M
- $0.075
- Context
- 128,000
- Max output
- 16,384
- Released
- 2024-07-18
Estimate your cost with GPT-4o-miniCompare all OpenAI modelsCount tokens for GPT-4o-mini
Pricing
GPT-4o-mini API pricing
| Price per 1M tokens | Input | Output |
|---|---|---|
| Standard | $0.15 | $0.60 |
| Cached input (read) | $0.075 | – |
| Batch API | $0.075 | $0.30 |
Output costs 4.0× the input price, so long replies drive the bill more than long prompts. If your requests share a long fixed prefix, such as a system prompt or tool definitions, caching cuts that part to 50% of the normal input price.
Examples
What GPT-4o-mini costs in practice
| Workload | Tokens in / out | Requests/day | Per request | Per month |
|---|---|---|---|---|
| Support chatbot50% cached | 1,500 / 400 | 1,000 | $0.000409 | $12.43 |
| RAG app20% cached | 6,000 / 500 | 500 | $0.00111 | $16.88 |
| Coding agent80% cached | 40,000 / 2,000 | 200 | $0.0048 | $29.20 |
| Document summariserbatch | 8,000 / 600 | 300 | $0.00078 | $7.12 |
Change any number in the cost calculator.
Context
Context window
GPT-4o-mini accepts up to 128,000 tokens per request, roughly 96,000 English words or 192 pages. That budget is shared between your prompt and the reply, and a single reply is capped at 16,384 tokens.
Filling the whole window with one prompt costs about $0.0192 at list price, so for repeated long-document work, retrieval (sending only the relevant parts) or prompt caching is usually much cheaper.
Features
Capabilities
- Image input
- Yes
- Tool / function calling
- Yes
- Reasoning mode
- No
- Prompt caching
- Yes
- Structured output
- Yes
- Open weights
- No
- Accepts
- text, image, file
- Knowledge cutoff
- 2023-10-31
As declared by the provider’s API. It shows a feature exists, not how well it works.
API
GPT-4o-mini model ID
- OpenRouter
openai/gpt-4o-mini- OpenAI API
gpt-4o-mini(alias)gpt-4o-mini-2024-07-18(pinned)
OpenAI lists GPT-4o-mini as current. Every GPT model ID, with aliases and snapshots.
OpenAI
Where GPT-4o-mini sits in OpenAI’s line-up
Ranked by blended price (three parts input to one part output, $0.2625 per 1M for GPT-4o-mini), it is number 3 of OpenAI’s 10 current models here, counting from the cheapest. One step down, GPT-6 Luna costs $0.20 blended (24% less). Unlike GPT-4o-mini, it reads up to 1,050,000 tokens (128,000 here), caps a reply at 128,000 tokens (16,384 here) and has a reasoning mode. One step up, GPT-5.4 Mini costs $1.6875 (6.4× as much). Unlike GPT-4o-mini, it reads up to 400,000 tokens (128,000 here), caps a reply at 128,000 tokens (16,384 here) and has a reasoning mode. See every OpenAI model and price
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| gpt-oss-120b | $0.037 | $0.17 | 131,072 | 2025-08-05 |
| GPT-6 Luna | $0.10 | $0.50 | 1,050,000 | 2026-09-22 |
| GPT-4o-mini (this page) | $0.15 | $0.60 | 128,000 | 2024-07-18 |
| GPT-5.4 Mini | $0.75 | $4.50 | 400,000 | 2026-03-17 |
| GPT-6.1 Sol | $2 | $10 | 1,050,000 | 2026-09-29 |
| GPT-4o | $2.50 | $10 | 128,000 | 2024-05-13 |
| GPT-5.6 Terra | $2 | $12 | 1,050,000 | 2026-07-09 |
| GPT-5.5 | $5 | $30 | 1,050,000 | 2026-04-24 |
| GPT-6 Astra | $10 | $50 | 1,050,000 | 2026-09-04 |
| GPT-5.5 Pro | $30 | $180 | 1,050,000 | 2026-04-24 |
OpenAI’s current models with a page here, cheapest first. Older and superseded models are listed on the provider page and in the comparison table.
Save
Cheaper alternatives
- Qwen3.5-9BQwen57% cheaper
- Seed 1.6 FlashByteDance Seed50% cheaper
- Gemma 4 31BGoogle42% cheaper
- GPT-6 LunaOpenAI24% cheaper
- Claude Haiku 5.5Anthropic24% cheaper
Current models from major providers, one per provider, that keep image input, tool calling, a usable context window and output limit. Compared by blended price.
Compare
Similarly priced models
- Mistral Small 4Mistral$0.15 / $0.60
- DeepSeek V4 Pro 0423DeepSeek$0.2088 / $0.4176
- Qwen3 Coder NextQwen$0.12 / $0.80
- GLM 5.3 FlashZ.ai$0.15 / $0.50
- Claude Haiku 5.5Anthropic$0.10 / $0.50
Current models from other major providers, closest blended price.
FAQ
GPT-4o-mini questions
How much does GPT-4o-mini cost?
GPT-4o-mini costs $0.15 per million input tokens and $0.60 per million output tokens, with cached input at $0.075. A typical chatbot reply (1,500 tokens in, 400 out) costs about $0.000465. Use the LLM cost calculator for your own workload.
What is GPT-4o-mini’s context window?
128,000 tokens, about 96,000 English words or 192 printed pages, shared between your prompt and the reply. A single reply can be up to 16,384 tokens.
Does GPT-4o-mini support prompt caching and batch requests?
Yes, cached input is billed at $0.075 per million tokens (50% of the normal input price). A batch API is available at $0.075 input and $0.30 output per million tokens.
What are cheaper alternatives to GPT-4o-mini?
With the same essentials (image input and tool calling), the cheapest options are Qwen3.5-9B (57% cheaper), Seed 1.6 Flash (50% cheaper), Gemma 4 31B (42% cheaper). Cheaper doesn’t mean equivalent, so test them on your own prompts.
What is the API model ID for GPT-4o-mini?
On OpenRouter it is openai/gpt-4o-mini. In OpenAI’s own API, use gpt-4o-mini; OpenAI lists the model as current. The GPT model ID list has every alias and snapshot. Pass the ID as the model value in each request.
Is GPT-4o-mini open source?
No, GPT-4o-mini is only available through OpenAI’s API and partners; its weights aren’t published.