Tokens & Costs
AI model comparisons, head to head
19 head-to-head comparisons of the AI models developers most often weigh against each other: prices, limits and features side by side, what real workloads cost on each, and which to pick for what.
Every figure is computed from prices updated daily, so the comparisons follow every price change. To compare any other models, use the model comparison table or the cost calculator.
6 comparisons
Flagship models
The main models from the big providers, head to head.
- Claude Sonnet 5.5 vs GPT-6.1 Sol$2/$10 vs $2/$10 per 1MAbout the same cost for a support chatbot
- Gemini 3.8 Flash vs GPT-6.1 Sol$0.75/$3.75 vs $2/$10 per 1MGemini 3.8 Flash is 2.6× cheaper for a support chatbot
- Claude Sonnet 5.5 vs Gemini 3.8 Flash$2/$10 vs $0.75/$3.75 per 1MGemini 3.8 Flash is 2.6× cheaper for a support chatbot
- GPT-6.1 Sol vs Grok 4.7$2/$10 vs $2/$6 per 1MGrok 4.7 is 1.3× cheaper for a support chatbot
- Claude Opus 5.5 vs GPT-6.1 Sol$4/$20 vs $2/$10 per 1MGPT-6.1 Sol is 2.0× cheaper for a support chatbot
- Mistral Large 4 vs DeepSeek V4 Pro$0.68/$2.09 vs $0.2088/$0.4176 per 1MDeepSeek V4 Pro is 4.2× cheaper for a support chatbot
3 comparisons
Fast, low-cost models
Small and mid-sized models for high-volume work.
- Gemini 3.8 Flash vs GPT-6 Luna$0.75/$3.75 vs $0.10/$0.50 per 1MGPT-6 Luna is 7.5× cheaper for a support chatbot
- Claude Haiku 5.5 vs GPT-6 Luna$0.10/$0.50 vs $0.10/$0.50 per 1MAbout the same cost for a support chatbot
- Claude Haiku 5.5 vs Gemini 3.8 Flash$0.10/$0.50 vs $0.75/$3.75 per 1MClaude Haiku 5.5 is 7.5× cheaper for a support chatbot
4 comparisons
Same provider, different tier
Is the bigger model worth the price?
- Claude Opus 5.5 vs Claude Sonnet 5.5$4/$20 vs $2/$10 per 1MClaude Sonnet 5.5 is 2.0× cheaper for a support chatbot
- Claude Fable 5.1 vs Claude Opus 5.5$10/$50 vs $4/$20 per 1MClaude Opus 5.5 is 2.5× cheaper for a support chatbot
- Claude Haiku 5.5 vs Claude Sonnet 5.5$0.10/$0.50 vs $2/$10 per 1MClaude Haiku 5.5 is 20× cheaper for a support chatbot
- Gemini 3.5 Flash Lite vs Gemini 3.8 Flash$0.30/$2.50 vs $0.75/$3.75 per 1MGemini 3.5 Flash Lite is 1.7× cheaper for a support chatbot
6 comparisons
Open-weight models
Models whose weights you can download and host anywhere.
- DeepSeek V4 Pro vs Claude Sonnet 5.5$0.2088/$0.4176 vs $2/$10 per 1MDeepSeek V4 Pro is 17× cheaper for a support chatbot
- DeepSeek V4.1 Flash vs Qwen3.8 Flash$0.30/$1.20 vs $0.15/$0.47 per 1MQwen3.8 Flash is 2.3× cheaper for a support chatbot
- Kimi K3 vs GLM 5.3$0.80/$10 vs $1.40/$4.40 per 1MGLM 5.3 is 1.7× cheaper for a support chatbot
- Kimi K3 vs DeepSeek V4 Pro$0.80/$10 vs $0.2088/$0.4176 per 1MDeepSeek V4 Pro is 15× cheaper for a support chatbot
- Llama 4 Maverick vs Llama 4 Scout$0.1875/$0.6525 vs $0.10/$0.30 per 1MLlama 4 Scout is 1.6× cheaper for a support chatbot
- gpt-oss-120b vs Llama 4 Maverick$0.037/$0.17 vs $0.1875/$0.6525 per 1Mgpt-oss-120b is 3.6× cheaper for a support chatbot
Prices updated 2026-10-11 · USD per 1M input/output tokens
Method
How these comparisons work
Each page puts two models’ list prices (for open-weight models, the price of OpenRouter’s top-ranked host), cached and batch prices, long-prompt surcharges, context window, output limit, accepted inputs and declared features in one table, then prices four workloads on both: a support chatbot, a RAG app, a coding agent and a batch document summariser. The workloads apply each provider’s own caching and batch rules, so the totals reflect how you’d actually be billed rather than the headline price per token.
The “which to pick” section turns those numbers into plain criteria: the lowest bill, very long prompts, long replies, images and other inputs, repeated prompts, batch jobs and self-hosting. A criterion goes to one model only when the data shows a real difference (more than 5% on prices, a clearly larger window), and we never claim one model gives better answers, because price data can’t show that.
Per-token prices also hide a second difference: tokenizers. The same English text can be noticeably more tokens on one model than another, so each page shows tokens per 1,000 words where the tokenizer is public or the provider publishes a figure. The gaps can be large: in the widest pair here, Claude Haiku 5.5 runs the same support chatbot (1,000 requests a day) for $8.59 a month against $169.57.
FAQ
Model comparison questions
How do you choose which models to compare?
We only build a comparison when people actually search for that pair: we check that several other sites already answer the exact “X vs Y” query. Both models also need a full model page in our data (prices, context window and release date). We don’t generate every possible pair, because most of them would say nothing useful.
Do these comparisons include benchmark scores?
No. Benchmarks vary by test, version and settings, and sites often disagree. We compare what we can verify: list prices, published limits, declared features and what real workloads cost. For quality, run both models on a sample of your own prompts; the cost tables tell you what that choice is worth.
Why can two models with the same price give different bills?
Each model counts tokens with its own tokenizer, so the same text can be a different number of tokens. Caching rules, batch discounts and long-prompt surcharges also differ. Each comparison shows tokens per 1,000 English words where we know it, and prices four workloads with each provider’s own rules. See the token counter to count your own text.
How often are the comparisons updated?
Every number is recalculated from our price data, which refreshes daily from the OpenRouter models API and the LiteLLM price list. The current data is from 2026-10-11. If either model in a pair leaves the data, its comparison page is removed rather than left showing stale numbers.
Can I compare models that aren’t listed here?
Yes. The AI model comparison table puts all 330+ models side by side with filters for provider, context window and features, and the LLM cost calculator prices your own workload on any of them, with the cheapest alternatives listed.