Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Model comparison

gpt-oss-120b vs Llama 4 Maverick: pricing, context and features compared

gpt-oss-120b costs $0.037 per million input tokens and $0.17 per million output tokens; Llama 4 Maverick costs $0.1875 and $0.6525. Here is how they compare on specs, on what real workloads cost, and on which to pick for what, from prices updated daily.

OpenAI’s and Meta’s open-weight models are the two big US-made options you can run yourself or rent from many hosts. They differ in what they accept and whether they reason before answering.

Input / 1M
$0.037
Output / 1M
$0.17
Context
131,072
Max output
Not published

Open weights: this is the price of OpenRouter’s top-ranked host, which may differ from OpenAI’s own API price.

Full gpt-oss-120b pricing and specs
Input / 1M
$0.1875
Output / 1M
$0.6525
Context
128,000
Max output
16,384

Open weights: this is the price of OpenRouter’s top-ranked host, which may differ from Meta’s own API price.

Full Llama 4 Maverick pricing and specs

Verdict

gpt-oss-120b or Llama 4 Maverick: which to pick

gpt-oss-120b has the lower list prices: 5.1× cheaper per input token and 3.8× per output token. On our data, gpt-oss-120b has the edge for the lowest bill, very long prompts, batch jobs and a reasoning mode, and Llama 4 Maverick for non-text inputs and repeated prompts. For scale, a support chatbot handling 1,000 requests a day costs about $3.76 a month on gpt-oss-120b and $13.36 on Llama 4 Maverick.

Lowest bill for typical workloadsPick gpt-oss-120b
gpt-oss-120b costs less on all 4 workloads below, 2.4× to 5.9× cheaper than Llama 4 Maverick at list prices.
Very long prompts (whole codebases, long documents)Pick gpt-oss-120b
Both read about 128,000 tokens per request. One 100,000-token prompt with a 2,000-token reply costs $0.00404 on gpt-oss-120b and $0.0201 on Llama 4 Maverick. Neither has a long-prompt surcharge in our data.
Long replies (reports, large code files)Either
Llama 4 Maverick caps a reply at 16,384 tokens; gpt-oss-120b publishes no separate cap, so we can’t say which allows longer replies.
Images, audio, video or files in the promptPick Llama 4 Maverick
Llama 4 Maverick also accepts images, which gpt-oss-120b doesn’t.
Repeated long prompts (prompt caching)Pick Llama 4 Maverick
Llama 4 Maverick bills cached input at $0.05 per 1M; gpt-oss-120b publishes no cached price, so repeated prefixes are billed in full.
Overnight batch jobsPick gpt-oss-120b
gpt-oss-120b has a batch price ($0.0296 in, $0.136 out per 1M); Llama 4 Maverick has none in our data.
Self-hosting or choosing your own hostEither
Both have open weights, so you can run either yourself or buy it from several hosts; prices here are typical hosted prices.
Step-by-step reasoningPick gpt-oss-120b
gpt-oss-120b declares a reasoning mode in its API; Llama 4 Maverick doesn’t.

Criteria use only list prices, published limits and the features each provider declares. We don’t rank answer quality; test both models on your own prompts before you commit.

Specs

gpt-oss-120b vs Llama 4 Maverick side by side

Specs of gpt-oss-120b and Llama 4 Maverick; the better value on each row is marked
Specgpt-oss-120bLlama 4 Maverick
ProviderOpenAIMeta
Input / 1M tokens$0.037 (better)$0.1875
Output / 1M tokens$0.17 (better)$0.6525
Cached input / 1MNot published$0.05 (better)
Batch in / out$0.0296 / $0.136 (better)No batch price
Long-prompt priceNoneNone
Context window131,072128,000
Max outputNot published16,384
Acceptstexttext, images
Image inputNoYes (better)
Tool callingYesYes
Reasoning modeYes (better)No
Prompt cachingNoYes (better)
Structured outputYesYes
Open weightsYesYes
Knowledge cutoff2024-06-302024-08-31
First listed2025-08-052025-04-05

Prices in US dollars per million tokens. Highlighted: the lower price or larger limit where the difference is over 5%. Features are as declared by each provider’s API; “First listed” is when our data source first listed the model. Long-prompt prices apply to the whole request once the prompt passes the threshold.

Costs

What gpt-oss-120b and Llama 4 Maverick cost for real workloads

WorkloadTokens in / outRequests/daygpt-oss-120b / monthLlama 4 Maverick / monthCheaper
Support chatbotCalculator: gpt-oss-120b · Llama 4 Maverick1,500 / 4001,000$3.76$13.3650% cachedgpt-oss-120b (3.6×)
RAG appCalculator: gpt-oss-120b · Llama 4 Maverick6,000 / 500500$4.67$19.5620% cachedgpt-oss-120b (4.2×)
Coding agentCalculator: gpt-oss-120b · Llama 4 Maverick40,000 / 2,000200$11.07$26.8080% cachedgpt-oss-120b (2.4×)
Document summariserCalculator: gpt-oss-120b · Llama 4 Maverick8,000 / 600300$2.91batch$17.26no batch pricegpt-oss-120b (5.9×)

Each row uses the same token counts for both models and list prices (for an open-weight model, the top-ranked host’s price on OpenRouter), with caching and batch discounts only where the provider publishes them. A month is 365 ÷ 12 days. The support chatbot reads 50% of its prompt from the cache. The RAG app reads 20% of its prompt from the cache. The coding agent reads 80% of its prompt from the cache. The document summariser runs as a batch job where a batch price exists. Open either model in the LLM cost calculator to change any number.

Tokenizers

Are their per-token prices comparable?

Not exactly. Each model splits text into tokens its own way, so the same prompt can be a different number of tokens on gpt-oss-120b and Llama 4 Maverick, and the cheaper per-token price isn’t always the cheaper bill. Where a tokenizer isn’t public or measured we say so rather than guess.

ModelTokenizerTokens per 1,000 wordsOutput per 1M words
gpt-oss-120bOpenAI o200k_baseOur measurement1,155$0.20
Llama 4 MaverickWe haven’t measured Llama 4 Maverick’s tokenizer, so its count for a given text isn’t known here.

English prose, measured on the Universal Declaration of Human Rights (2026-10-11) or taken from the provider’s published words-per-token figure. Code, JSON and other languages use more tokens per word. Count your own text in the token counter or convert with tokens to words.

FAQ

gpt-oss-120b vs Llama 4 Maverick questions

Is gpt-oss-120b cheaper than Llama 4 Maverick?

gpt-oss-120b costs $0.037 per million input tokens and $0.17 per million output tokens; Llama 4 Maverick costs $0.1875 and $0.6525. For a support chatbot handling 1,000 requests a day, that is about $3.76 a month on gpt-oss-120b against $13.36 on Llama 4 Maverick, so gpt-oss-120b is 3.6× cheaper there. Try your own numbers in the LLM cost calculator.

Which has the bigger context window, gpt-oss-120b or Llama 4 Maverick?

Their context windows are about the same size. gpt-oss-120b reads up to 131,072 tokens (roughly 98,304 English words) and publishes no separate cap on reply length. Llama 4 Maverick reads up to 128,000 tokens (roughly 96,000 English words) and writes up to 16,384 tokens per reply. The window is shared between your prompt and the reply. Word counts use a rough 0.75 words per token; real ratios depend on the tokenizer.

Is gpt-oss-120b or Llama 4 Maverick better for coding?

We don’t publish benchmark scores, so this page can’t say which writes better code. What we can show is cost: a coding agent re-sending 40,000 tokens of context (80% cached) and writing 2,000 tokens, 200 times a day, costs about $11.07 a month on gpt-oss-120b and $26.80 on Llama 4 Maverick. Both support tool calling. Test both on tasks from your own repository.

Can gpt-oss-120b and Llama 4 Maverick read images?

gpt-oss-120b accepts text only. Llama 4 Maverick accepts images as well as text. Only Llama 4 Maverick can read screenshots, photos or scanned pages, so for image work the choice is made for you. This is what each provider declares for its API, not a measure of how well it works.

Are gpt-oss-120b and Llama 4 Maverick open source?

Both have open weights: you can download them, run them on your own hardware, or buy them from several hosting providers. The prices on this page are typical hosted prices, so shop around. Check each model’s licence for commercial use, and see the VRAM calculator for the hardware a model needs.

Do gpt-oss-120b and Llama 4 Maverick count tokens the same way?

Not necessarily. Each model splits text into tokens with its own tokenizer, so the same prompt can be a different number of tokens on each, and per-token prices aren’t directly comparable. gpt-oss-120b uses about 1,155 tokens per 1,000 English words (our measurement). We haven’t measured Llama 4 Maverick’s tokenizer, so its count for a given text isn’t known here. Count your own text with the token counter or convert with tokens to words.

Updated

Sources: OpenRouter models API and LiteLLM model prices and context windows, checked daily; open-weight prices are typical hosted prices. Tokenizer figures: our measurements (see tokens to words). Confirm critical numbers on each provider’s pricing page. See our methodology.