Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Meta · Model

Llama 4 Maverick pricing, context window and specs

Llama 4 Maverick costs $0.1875 per million input tokens and $0.6525 per million output tokens, and reads up to 128,000 tokens per request. By blended price it is number 9 of 14 from cheapest in Meta's current line-up.

Input / 1M
$0.1875
Output / 1M
$0.6525
Cached / 1M
$0.05
Context
128,000
Max output
16,384
Released
2025-04-05

Estimate your cost with Llama 4 MaverickCompare all Meta modelsCount tokens for Llama 4 Maverick

Pricing

Llama 4 Maverick API pricing

Price per 1M tokensInputOutput
Standard$0.1875$0.6525
Cached input (read)$0.05–

Output costs 3.5× the input price, so long replies drive the bill more than long prompts. If your requests share a long fixed prefix, such as a system prompt or tool definitions, caching cuts that part to 27% of the normal input price.

Examples

What Llama 4 Maverick costs in practice

WorkloadTokens in / outRequests/dayPer requestPer month
Support chatbot50% cached1,500 / 4001,000$0.000439$13.36
RAG app20% cached6,000 / 500500$0.00129$19.56
Coding agent80% cached40,000 / 2,000200$0.0044$26.80
Document summariser8,000 / 600300$0.00189$17.26

Change any number in the cost calculator.

Context

Context window

Llama 4 Maverick accepts up to 128,000 tokens per request, roughly 96,000 English words or 192 pages. That budget is shared between your prompt and the reply, and a single reply is capped at 16,384 tokens.

Filling the whole window with one prompt costs about $0.024 at list price, so for repeated long-document work, retrieval (sending only the relevant parts) or prompt caching is usually much cheaper.

Check whether your text fits Llama 4 Maverick’s window.

Features

Capabilities

Image input
Yes
Tool / function calling
Yes
Reasoning mode
No
Prompt caching
Yes
Structured output
Yes
Open weights
Yes
Accepts
text, image
Knowledge cutoff
2024-08-31

As declared by the provider’s API. It shows a feature exists, not how well it works.

Meta

Where Llama 4 Maverick sits in Meta’s line-up

Ranked by blended price (three parts input to one part output, $0.3037 per 1M for Llama 4 Maverick), it is number 9 of 14 from cheapest of Meta’s 14 priced models. See every Meta model and price

ModelInput / 1MOutput / 1MContextReleased
Muse Spark 1.3$1.25$4.251,048,5762026-09-02
Muse Spark 1.3 Contributor$0.10$0.201,048,5762026-09-02
Muse Spark 1.2 Contributor$0.10$0.201,048,5762026-08-21
Muse Glimmer 30B$0.30$1.20131,0722026-08-09
Muse Spark 1.2$1.25$4.251,048,5762026-08-05
Muse Spark 1.1$1.25$4.251,048,5762026-07-16
Llama Guard 4 12B$0.18$0.18163,8402025-04-30
Llama 4 Maverick (this page)$0.1875$0.6525128,0002025-04-05

Save

Cheaper alternatives

  • Qwen3.7 FlashQwen82% cheaper
  • Seed 1.6 FlashByteDance Seed57% cheaper
  • Gemma 4 31BGoogle50% cheaper
  • Claude Haiku 5.5Anthropic34% cheaper
  • GPT-6 LunaOpenAI34% cheaper

Current models from major providers, one per provider, that keep image input, tool calling, a usable context window and output limit. Compared by blended price.

Compare

Similarly priced models

  • DeepSeek V3.2DeepSeek$0.259 / $0.42
  • Qwen3 Coder NextQwen$0.12 / $0.80
  • Mistral Small 4Mistral$0.15 / $0.60
  • MiniMax M2.7MiniMax$0.21 / $0.84
  • GLM 5.3 FlashZ.ai$0.15 / $0.50

Current models from other major providers, closest blended price.

FAQ

Llama 4 Maverick questions

How much does Llama 4 Maverick cost?

Llama 4 Maverick costs $0.1875 per million input tokens and $0.6525 per million output tokens, with cached input at $0.05. A typical chatbot reply (1,500 tokens in, 400 out) costs about $0.000542. Use the LLM cost calculator for your own workload.

What is Llama 4 Maverick's context window?

128,000 tokens, about 96,000 English words or 192 printed pages, shared between your prompt and the reply. A single reply can be up to 16,384 tokens.

Does Llama 4 Maverick support prompt caching and batch requests?

Yes, cached input is billed at $0.05 per million tokens (27% of the normal input price). No batch price is published for it.

What are cheaper alternatives to Llama 4 Maverick?

With the same essentials (image input and tool calling), the cheapest options are Qwen3.7 Flash (82% cheaper), Seed 1.6 Flash (57% cheaper), Gemma 4 31B (50% cheaper). Cheaper doesn't mean equivalent, so test them on your own prompts.

Is Llama 4 Maverick open source?

Yes, Llama 4 Maverick has open weights, so you can run it yourself or choose between several hosting providers. The prices here are a typical hosted price.

Updated

Sources: OpenRouter models API and LiteLLM model prices and context windows, checked daily; open-weight prices are typical hosted prices. Confirm critical numbers on Meta’s pricing page. See our methodology.