Meta · Model
Llama 4 Maverick pricing, context window and specs
Llama 4 Maverick costs $0.1875 per million input tokens and $0.6525 per million output tokens, and reads up to 128,000 tokens per request. By blended price it is number 9 of 14 from cheapest in Meta's current line-up.
- Input / 1M
- $0.1875
- Output / 1M
- $0.6525
- Cached / 1M
- $0.05
- Context
- 128,000
- Max output
- 16,384
- Released
- 2025-04-05
Estimate your cost with Llama 4 MaverickCompare all Meta modelsCount tokens for Llama 4 Maverick
Pricing
Llama 4 Maverick API pricing
| Price per 1M tokens | Input | Output |
|---|---|---|
| Standard | $0.1875 | $0.6525 |
| Cached input (read) | $0.05 | – |
Output costs 3.5× the input price, so long replies drive the bill more than long prompts. If your requests share a long fixed prefix, such as a system prompt or tool definitions, caching cuts that part to 27% of the normal input price.
Examples
What Llama 4 Maverick costs in practice
| Workload | Tokens in / out | Requests/day | Per request | Per month |
|---|---|---|---|---|
| Support chatbot50% cached | 1,500 / 400 | 1,000 | $0.000439 | $13.36 |
| RAG app20% cached | 6,000 / 500 | 500 | $0.00129 | $19.56 |
| Coding agent80% cached | 40,000 / 2,000 | 200 | $0.0044 | $26.80 |
| Document summariser | 8,000 / 600 | 300 | $0.00189 | $17.26 |
Change any number in the cost calculator.
Context
Context window
Llama 4 Maverick accepts up to 128,000 tokens per request, roughly 96,000 English words or 192 pages. That budget is shared between your prompt and the reply, and a single reply is capped at 16,384 tokens.
Filling the whole window with one prompt costs about $0.024 at list price, so for repeated long-document work, retrieval (sending only the relevant parts) or prompt caching is usually much cheaper.
Features
Capabilities
- Image input
- Yes
- Tool / function calling
- Yes
- Reasoning mode
- No
- Prompt caching
- Yes
- Structured output
- Yes
- Open weights
- Yes
- Accepts
- text, image
- Knowledge cutoff
- 2024-08-31
As declared by the provider’s API. It shows a feature exists, not how well it works.
Meta
Where Llama 4 Maverick sits in Meta’s line-up
Ranked by blended price (three parts input to one part output, $0.3037 per 1M for Llama 4 Maverick), it is number 9 of 14 from cheapest of Meta’s 14 priced models. See every Meta model and price
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| Muse Spark 1.3 | $1.25 | $4.25 | 1,048,576 | 2026-09-02 |
| Muse Spark 1.3 Contributor | $0.10 | $0.20 | 1,048,576 | 2026-09-02 |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 | 1,048,576 | 2026-08-21 |
| Muse Glimmer 30B | $0.30 | $1.20 | 131,072 | 2026-08-09 |
| Muse Spark 1.2 | $1.25 | $4.25 | 1,048,576 | 2026-08-05 |
| Muse Spark 1.1 | $1.25 | $4.25 | 1,048,576 | 2026-07-16 |
| Llama Guard 4 12B | $0.18 | $0.18 | 163,840 | 2025-04-30 |
| Llama 4 Maverick (this page) | $0.1875 | $0.6525 | 128,000 | 2025-04-05 |
Save
Cheaper alternatives
- Qwen3.7 FlashQwen82% cheaper
- Seed 1.6 FlashByteDance Seed57% cheaper
- Gemma 4 31BGoogle50% cheaper
- Claude Haiku 5.5Anthropic34% cheaper
- GPT-6 LunaOpenAI34% cheaper
Current models from major providers, one per provider, that keep image input, tool calling, a usable context window and output limit. Compared by blended price.
Compare
Similarly priced models
- DeepSeek V3.2DeepSeek$0.259 / $0.42
- Qwen3 Coder NextQwen$0.12 / $0.80
- Mistral Small 4Mistral$0.15 / $0.60
- MiniMax M2.7MiniMax$0.21 / $0.84
- GLM 5.3 FlashZ.ai$0.15 / $0.50
Current models from other major providers, closest blended price.
FAQ
Llama 4 Maverick questions
How much does Llama 4 Maverick cost?
Llama 4 Maverick costs $0.1875 per million input tokens and $0.6525 per million output tokens, with cached input at $0.05. A typical chatbot reply (1,500 tokens in, 400 out) costs about $0.000542. Use the LLM cost calculator for your own workload.
What is Llama 4 Maverick's context window?
128,000 tokens, about 96,000 English words or 192 printed pages, shared between your prompt and the reply. A single reply can be up to 16,384 tokens.
Does Llama 4 Maverick support prompt caching and batch requests?
Yes, cached input is billed at $0.05 per million tokens (27% of the normal input price). No batch price is published for it.
What are cheaper alternatives to Llama 4 Maverick?
With the same essentials (image input and tool calling), the cheapest options are Qwen3.7 Flash (82% cheaper), Seed 1.6 Flash (57% cheaper), Gemma 4 31B (50% cheaper). Cheaper doesn't mean equivalent, so test them on your own prompts.
Is Llama 4 Maverick open source?
Yes, Llama 4 Maverick has open weights, so you can run it yourself or choose between several hosting providers. The prices here are a typical hosted price.