Model comparison
gpt-oss-120b vs Llama 4 Maverick: pricing, context and features compared
gpt-oss-120b costs $0.037 per million input tokens and $0.17 per million output tokens; Llama 4 Maverick costs $0.1875 and $0.6525. Here is how they compare on specs, on what real workloads cost, and on which to pick for what, from prices updated daily.
OpenAI’s and Meta’s open-weight models are the two big US-made options you can run yourself or rent from many hosts. They differ in what they accept and whether they reason before answering.
- Input / 1M
- $0.037
- Output / 1M
- $0.17
- Context
- 131,072
- Max output
- Not published
Open weights: this is the price of OpenRouter’s top-ranked host, which may differ from OpenAI’s own API price.
Full gpt-oss-120b pricing and specs- Input / 1M
- $0.1875
- Output / 1M
- $0.6525
- Context
- 128,000
- Max output
- 16,384
Open weights: this is the price of OpenRouter’s top-ranked host, which may differ from Meta’s own API price.
Full Llama 4 Maverick pricing and specsVerdict
gpt-oss-120b or Llama 4 Maverick: which to pick
gpt-oss-120b has the lower list prices: 5.1× cheaper per input token and 3.8× per output token. On our data, gpt-oss-120b has the edge for the lowest bill, very long prompts, batch jobs and a reasoning mode, and Llama 4 Maverick for non-text inputs and repeated prompts. For scale, a support chatbot handling 1,000 requests a day costs about $3.76 a month on gpt-oss-120b and $13.36 on Llama 4 Maverick.
- Lowest bill for typical workloadsPick gpt-oss-120b
- gpt-oss-120b costs less on all 4 workloads below, 2.4× to 5.9× cheaper than Llama 4 Maverick at list prices.
- Very long prompts (whole codebases, long documents)Pick gpt-oss-120b
- Both read about 128,000 tokens per request. One 100,000-token prompt with a 2,000-token reply costs $0.00404 on gpt-oss-120b and $0.0201 on Llama 4 Maverick. Neither has a long-prompt surcharge in our data.
- Long replies (reports, large code files)Either
- Llama 4 Maverick caps a reply at 16,384 tokens; gpt-oss-120b publishes no separate cap, so we can’t say which allows longer replies.
- Images, audio, video or files in the promptPick Llama 4 Maverick
- Llama 4 Maverick also accepts images, which gpt-oss-120b doesn’t.
- Repeated long prompts (prompt caching)Pick Llama 4 Maverick
- Llama 4 Maverick bills cached input at $0.05 per 1M; gpt-oss-120b publishes no cached price, so repeated prefixes are billed in full.
- Overnight batch jobsPick gpt-oss-120b
- gpt-oss-120b has a batch price ($0.0296 in, $0.136 out per 1M); Llama 4 Maverick has none in our data.
- Self-hosting or choosing your own hostEither
- Both have open weights, so you can run either yourself or buy it from several hosts; prices here are typical hosted prices.
- Step-by-step reasoningPick gpt-oss-120b
- gpt-oss-120b declares a reasoning mode in its API; Llama 4 Maverick doesn’t.
Criteria use only list prices, published limits and the features each provider declares. We don’t rank answer quality; test both models on your own prompts before you commit.
Specs
gpt-oss-120b vs Llama 4 Maverick side by side
| Spec | gpt-oss-120b | Llama 4 Maverick |
|---|---|---|
| Provider | OpenAI | Meta |
| Input / 1M tokens | $0.037 (better) | $0.1875 |
| Output / 1M tokens | $0.17 (better) | $0.6525 |
| Cached input / 1M | Not published | $0.05 (better) |
| Batch in / out | $0.0296 / $0.136 (better) | No batch price |
| Long-prompt price | None | None |
| Context window | 131,072 | 128,000 |
| Max output | Not published | 16,384 |
| Accepts | text | text, images |
| Image input | No | Yes (better) |
| Tool calling | Yes | Yes |
| Reasoning mode | Yes (better) | No |
| Prompt caching | No | Yes (better) |
| Structured output | Yes | Yes |
| Open weights | Yes | Yes |
| Knowledge cutoff | 2024-06-30 | 2024-08-31 |
| First listed | 2025-08-05 | 2025-04-05 |
Prices in US dollars per million tokens. Highlighted: the lower price or larger limit where the difference is over 5%. Features are as declared by each provider’s API; “First listed” is when our data source first listed the model. Long-prompt prices apply to the whole request once the prompt passes the threshold.
Costs
What gpt-oss-120b and Llama 4 Maverick cost for real workloads
| Workload | Tokens in / out | Requests/day | gpt-oss-120b / month | Llama 4 Maverick / month | Cheaper |
|---|---|---|---|---|---|
| Support chatbotCalculator: gpt-oss-120b · Llama 4 Maverick | 1,500 / 400 | 1,000 | $3.76 | $13.3650% cached | gpt-oss-120b (3.6×) |
| RAG appCalculator: gpt-oss-120b · Llama 4 Maverick | 6,000 / 500 | 500 | $4.67 | $19.5620% cached | gpt-oss-120b (4.2×) |
| Coding agentCalculator: gpt-oss-120b · Llama 4 Maverick | 40,000 / 2,000 | 200 | $11.07 | $26.8080% cached | gpt-oss-120b (2.4×) |
| Document summariserCalculator: gpt-oss-120b · Llama 4 Maverick | 8,000 / 600 | 300 | $2.91batch | $17.26no batch price | gpt-oss-120b (5.9×) |
Each row uses the same token counts for both models and list prices (for an open-weight model, the top-ranked host’s price on OpenRouter), with caching and batch discounts only where the provider publishes them. A month is 365 ÷ 12 days. The support chatbot reads 50% of its prompt from the cache. The RAG app reads 20% of its prompt from the cache. The coding agent reads 80% of its prompt from the cache. The document summariser runs as a batch job where a batch price exists. Open either model in the LLM cost calculator to change any number.
Tokenizers
Are their per-token prices comparable?
Not exactly. Each model splits text into tokens its own way, so the same prompt can be a different number of tokens on gpt-oss-120b and Llama 4 Maverick, and the cheaper per-token price isn’t always the cheaper bill. Where a tokenizer isn’t public or measured we say so rather than guess.
| Model | Tokenizer | Tokens per 1,000 words | Output per 1M words |
|---|---|---|---|
| gpt-oss-120b | OpenAI o200k_baseOur measurement | 1,155 | $0.20 |
| Llama 4 Maverick | We haven’t measured Llama 4 Maverick’s tokenizer, so its count for a given text isn’t known here. | ||
English prose, measured on the Universal Declaration of Human Rights (2026-10-11) or taken from the provider’s published words-per-token figure. Code, JSON and other languages use more tokens per word. Count your own text in the token counter or convert with tokens to words.
FAQ
gpt-oss-120b vs Llama 4 Maverick questions
Is gpt-oss-120b cheaper than Llama 4 Maverick?
gpt-oss-120b costs $0.037 per million input tokens and $0.17 per million output tokens; Llama 4 Maverick costs $0.1875 and $0.6525. For a support chatbot handling 1,000 requests a day, that is about $3.76 a month on gpt-oss-120b against $13.36 on Llama 4 Maverick, so gpt-oss-120b is 3.6× cheaper there. Try your own numbers in the LLM cost calculator.
Which has the bigger context window, gpt-oss-120b or Llama 4 Maverick?
Their context windows are about the same size. gpt-oss-120b reads up to 131,072 tokens (roughly 98,304 English words) and publishes no separate cap on reply length. Llama 4 Maverick reads up to 128,000 tokens (roughly 96,000 English words) and writes up to 16,384 tokens per reply. The window is shared between your prompt and the reply. Word counts use a rough 0.75 words per token; real ratios depend on the tokenizer.
Is gpt-oss-120b or Llama 4 Maverick better for coding?
We don’t publish benchmark scores, so this page can’t say which writes better code. What we can show is cost: a coding agent re-sending 40,000 tokens of context (80% cached) and writing 2,000 tokens, 200 times a day, costs about $11.07 a month on gpt-oss-120b and $26.80 on Llama 4 Maverick. Both support tool calling. Test both on tasks from your own repository.
Can gpt-oss-120b and Llama 4 Maverick read images?
gpt-oss-120b accepts text only. Llama 4 Maverick accepts images as well as text. Only Llama 4 Maverick can read screenshots, photos or scanned pages, so for image work the choice is made for you. This is what each provider declares for its API, not a measure of how well it works.
Are gpt-oss-120b and Llama 4 Maverick open source?
Both have open weights: you can download them, run them on your own hardware, or buy them from several hosting providers. The prices on this page are typical hosted prices, so shop around. Check each model’s licence for commercial use, and see the VRAM calculator for the hardware a model needs.
Do gpt-oss-120b and Llama 4 Maverick count tokens the same way?
Not necessarily. Each model splits text into tokens with its own tokenizer, so the same prompt can be a different number of tokens on each, and per-token prices aren’t directly comparable. gpt-oss-120b uses about 1,155 tokens per 1,000 English words (our measurement). We haven’t measured Llama 4 Maverick’s tokenizer, so its count for a given text isn’t known here. Count your own text with the token counter or convert with tokens to words.
Explore
Go further
- gpt-oss-120b pricing, context window and alternatives
- Llama 4 Maverick pricing, context window and alternatives
- Every OpenAI and Meta model in the comparison table
- OpenAI API pricing for every model
- Meta API pricing for every model
Guides
Compare