Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Meta · Model

Llama 4 Scout pricing, context window and specs

Llama 4 Scout costs $0.10 per million input tokens and $0.30 per million output tokens, and reads up to 327,680 tokens per request. By blended price it is number 6 of 14 from cheapest in Meta's current line-up.

Input / 1M
$0.10
Output / 1M
$0.30
Cached / 1M
–
Context
327,680
Max output
16,384
Released
2025-04-05

Estimate your cost with Llama 4 ScoutCompare all Meta modelsCount tokens for Llama 4 Scout

Pricing

Llama 4 Scout API pricing

Price per 1M tokensInputOutput
Standard$0.10$0.30

Output costs 3.0× the input price, so long replies drive the bill more than long prompts. No cached-input price is published, so repeated prompt prefixes are billed at the full input price.

Examples

What Llama 4 Scout costs in practice

WorkloadTokens in / outRequests/dayPer requestPer month
Support chatbot1,500 / 4001,000$0.00027$8.21
RAG app6,000 / 500500$0.00075$11.41
Coding agent40,000 / 2,000200$0.0046$27.98
Document summariser8,000 / 600300$0.00098$8.94

Change any number in the cost calculator.

Context

Context window

Llama 4 Scout accepts up to 327,680 tokens per request, roughly 245,760 English words or 492 pages. That budget is shared between your prompt and the reply, and a single reply is capped at 16,384 tokens.

Filling the whole window with one prompt costs about $0.0328 at list price, so for repeated long-document work, retrieval (sending only the relevant parts) or prompt caching is usually much cheaper.

Check whether your text fits Llama 4 Scout’s window.

Features

Capabilities

Image input
Yes
Tool / function calling
Yes
Reasoning mode
No
Prompt caching
No
Structured output
Yes
Open weights
Yes
Accepts
text, image
Knowledge cutoff
2024-08-31

As declared by the provider’s API. It shows a feature exists, not how well it works.

Meta

Where Llama 4 Scout sits in Meta’s line-up

Ranked by blended price (three parts input to one part output, $0.15 per 1M for Llama 4 Scout), it is number 6 of 14 from cheapest of Meta’s 14 priced models. See every Meta model and price

ModelInput / 1MOutput / 1MContextReleased
Muse Spark 1.3$1.25$4.251,048,5762026-09-02
Muse Spark 1.3 Contributor$0.10$0.201,048,5762026-09-02
Muse Spark 1.2 Contributor$0.10$0.201,048,5762026-08-21
Muse Glimmer 30B$0.30$1.20131,0722026-08-09
Muse Spark 1.2$1.25$4.251,048,5762026-08-05
Muse Spark 1.1$1.25$4.251,048,5762026-07-16
Llama Guard 4 12B$0.18$0.18163,8402025-04-30
Llama 4 Maverick$0.1875$0.6525128,0002025-04-05

Save

Cheaper alternatives

  • Qwen3.7 FlashQwen63% cheaper
  • Seed 1.6 FlashByteDance Seed13% cheaper

Current models from major providers, one per provider, that keep image input, tool calling, a usable context window and output limit. Compared by blended price.

Compare

Similarly priced models

  • Ministral 3 8B 2512Mistral$0.15 / $0.15
  • Gemma 4 31BGoogle$0.09 / $0.34
  • GLM 4.7 FlashZ.ai$0.0605 / $0.40
  • Seed 1.6 FlashByteDance Seed$0.075 / $0.30
  • Nemotron 3 SuperNVIDIA$0.08 / $0.45

Current models from other major providers, closest blended price.

FAQ

Llama 4 Scout questions

How much does Llama 4 Scout cost?

Llama 4 Scout costs $0.10 per million input tokens and $0.30 per million output tokens. A typical chatbot reply (1,500 tokens in, 400 out) costs about $0.00027. Use the LLM cost calculator for your own workload.

What is Llama 4 Scout's context window?

327,680 tokens, about 245,760 English words or 492 printed pages, shared between your prompt and the reply. A single reply can be up to 16,384 tokens.

Does Llama 4 Scout support prompt caching and batch requests?

No cached-input price is published for it. No batch price is published for it.

What are cheaper alternatives to Llama 4 Scout?

With the same essentials (image input and tool calling), the cheapest options are Qwen3.7 Flash (63% cheaper), Seed 1.6 Flash (13% cheaper). Cheaper doesn't mean equivalent, so test them on your own prompts.

Is Llama 4 Scout open source?

Yes, Llama 4 Scout has open weights, so you can run it yourself or choose between several hosting providers. The prices here are a typical hosted price.

Updated

Sources: OpenRouter models API and LiteLLM model prices and context windows, checked daily; open-weight prices are typical hosted prices. Confirm critical numbers on Meta’s pricing page. See our methodology.