01 — API Pricing Index

Every LLM price,
on one honest scale.

Input, output, and cached-input rates for the models developers actually ship on — normalized to USD per million tokens, so comparison is finally apples-to-apples.

Project my monthly bill →How it works
Index SummaryUSD / 1M
Models tracked
108
across 9 providers
Providers
9
updated 2026-08-27
Cheapest input
$0.03
Qwen3.7 Flash
Priciest output
$180.00
GPT-5.4 Pro
New Model Releases
View all →
Qwen3.8 FlashAug 26
GLM 5.3 FlashAug 26
DeepSeek V4 Flash Vision ExpAug 21
GLM 5.3Aug 18
Qwen3.8 27BAug 14
Gemini 3.7 Flash (batch)Aug 13
Gemini 3.7 FlashAug 13
Qwen3.8 2.4T A95BAug 12
DeepSeek V4 Pro 0813Aug 12
Grok 4.6Aug 12
Qwen3.8 MaxAug 3
DeepSeek V4 Flash 0731Jul 31
Price Trends
287 price cuts, 223 increases logged so far
A running, dated log of exactly how the market is moving — not just a snapshot.
See price trends →
02 — The Table

All models, sorted your way

YEAR
SORT
WEIGHTSSee the open-vs-closed race →
ModelProviderReleased ↓ContextInputOutputCached inSpeedTTFT
QwenQwen3.7 FlashBudget
QwenJul 20261M$0.03$0.13$0.01
Qwen · 1M context · Jul 2026 · cached $0.01
DeepSeekJul 20261M$0.05$0.10$0.01
DeepSeek · 1M context · Jul 2026 · cached $0.01 · open-weight
GLM 4.7 FlashBudgetOPEN
Z.AIJan 2026203K$0.06$0.40$0.0190.1 t/s1.38s
Z.AI · 203K context · Jan 2026 · cached $0.01 · open-weight · 90.1 t/s · 1.38s TTFT
QwenQwen3.5-FlashBudget
QwenFeb 20261M$0.07$0.26$0.01
Qwen · 1M context · Feb 2026 · cached $0.01
GoogleGemma 4 26B A4B BudgetOPEN
GoogleApr 2026262K$0.07$0.34$0.01
Google · 262K context · Apr 2026 · cached $0.01 · open-weight
GLM 5.3 FlashBudgetOPEN
Z.AIAug 20261M$0.07$0.25$0.01
Z.AI · 1M context · Aug 2026 · cached $0.01 · open-weight
DeepSeekApr 20261M$0.08$0.16$0.02
DeepSeek · 1M context · Apr 2026 · cached $0.02 · open-weight
GoogleGemma 4 31BBudgetOPEN
GoogleApr 2026262K$0.09$0.34$0.05
Google · 262K context · Apr 2026 · cached $0.05 · open-weight
OpenAIMar 2026400K$0.10$0.63$0.01
OpenAI · 400K context · Mar 2026 · cached $0.01
OpenAIJul 20261M$0.10$0.60$0.01
OpenAI · 1M context · Jul 2026 · cached $0.01
OpenAIJul 20261M$0.10$0.60$0.01
OpenAI · 1M context · Jul 2026 · cached $0.01
QwenQwen3.5-9BBudgetOPEN
QwenMar 2026262K$0.10$0.15$0.0176.7 t/s1.62s
Qwen · 262K context · Mar 2026 · cached $0.01 · open-weight · 76.7 t/s · 1.62s TTFT
QwenQwen3.6 35B A3BBudgetOPEN
QwenApr 2026262K$0.10$0.90$0.05
Qwen · 262K context · Apr 2026 · cached $0.05 · open-weight
QwenQwen3 Coder NextBudgetOPEN
QwenFeb 2026262K$0.12$0.80$0.07
Qwen · 262K context · Feb 2026 · cached $0.07 · open-weight
GoogleMay 20261M$0.13$0.75$0.01
Google · 1M context · May 2026 · cached $0.01
GoogleJul 20261M$0.15$1.25$0.01
Google · 1M context · Jul 2026 · cached $0.01
MistralMistral Small 4BudgetOPEN
MistralMar 2026262K$0.15$0.60$0.01176.7 t/s0.75s
Mistral · 262K context · Mar 2026 · cached $0.01 · open-weight · 176.7 t/s · 0.75s TTFT
QwenQwen3.8 FlashBudget
QwenAug 20261M$0.15$0.47$0.02
Qwen · 1M context · Aug 2026 · cached $0.02
GoogleAug 20261M$0.19$0.94$0.02
Google · 1M context · Aug 2026 · cached $0.02
QwenQwen3.6 FlashBudget
QwenApr 20261M$0.19$1.13$0.02
Qwen · 1M context · Apr 2026 · cached $0.02
QwenQwen3.5-27BBudgetOPEN
QwenFeb 2026262K$0.20$1.56$0.0278.3 t/s5.61s
Qwen · 262K context · Feb 2026 · cached $0.02 · open-weight · 78.3 t/s · 5.61s TTFT
OpenAIGPT-5.4 NanoBudget
OpenAIMar 2026400K$0.20$1.25$0.02
OpenAI · 400K context · Mar 2026 · cached $0.02
OpenAIGPT-5.6 LunaBudget
OpenAIJul 20261M$0.20$1.20$0.02130.9 t/s134.6s
OpenAI · 1M context · Jul 2026 · cached $0.02 · 130.9 t/s · 134.6s TTFT
OpenAIJul 20261M$0.20$1.20$0.02
OpenAI · 1M context · Jul 2026 · cached $0.02
DeepSeekAug 20261M$0.22$0.66$0.01
DeepSeek · 1M context · Aug 2026 · cached $0.01
GoogleMay 20261M$0.25$1.50$0.03
Google · 1M context · May 2026 · cached $0.03
GoogleMar 20261M$0.25$1.50$0.03
Google · 1M context · Mar 2026 · cached $0.03
QwenQwen3.5-35B-A3BBudgetOPEN
QwenFeb 2026262K$0.25$1.25$0.25153.1 t/s2.04s
Qwen · 262K context · Feb 2026 · cached $0.25 · open-weight · 153.1 t/s · 2.04s TTFT
QwenFeb 20261M$0.26$1.56$0.03
Qwen · 1M context · Feb 2026 · cached $0.03
QwenQwen3.5-122B-A10BBudgetOPEN
QwenFeb 2026262K$0.26$2.08$0.03135.2 t/s2.3s
Qwen · 262K context · Feb 2026 · cached $0.03 · open-weight · 135.2 t/s · 2.3s TTFT
GoogleJul 20261M$0.30$2.50$0.03412.6 t/s8.6s
Google · 1M context · Jul 2026 · cached $0.03 · 412.6 t/s · 8.6s TTFT
QwenApr 20261M$0.30$1.80$0.03
Qwen · 1M context · Apr 2026 · cached $0.03
QwenQwen3.7 PlusBudget
QwenJun 20261M$0.32$1.28$0.06
Qwen · 1M context · Jun 2026 · cached $0.06
QwenQwen3.6 PlusBudget
QwenApr 20261M$0.33$1.95$0.03
Qwen · 1M context · Apr 2026 · cached $0.03
GoogleJul 20261M$0.38$1.88$0.04
Google · 1M context · Jul 2026 · cached $0.04
GoogleAug 20261M$0.38$1.88$0.04
Google · 1M context · Aug 2026 · cached $0.04
OpenAIMar 2026400K$0.38$2.25$0.04
OpenAI · 400K context · Mar 2026 · cached $0.04
QwenQwen3.5 397B A17BBudgetOPEN
QwenFeb 2026262K$0.39$2.34$0.0481.8 t/s2.19s
Qwen · 262K context · Feb 2026 · cached $0.04 · open-weight · 81.8 t/s · 2.19s TTFT
QwenQwen3.8 27BBudgetOPEN
QwenAug 20261M$0.42$2.55$0.09
Qwen · 1M context · Aug 2026 · cached $0.09 · open-weight
MoonshotKimi K2.5BudgetOPEN
MoonshotJan 2026262K$0.60$3.00$0.10
Moonshot · 262K context · Jan 2026 · cached $0.10 · open-weight
QwenQwen3.6 27BBudgetOPEN
QwenApr 2026262K$0.60$3.60$0.12
Qwen · 262K context · Apr 2026 · cached $0.12 · open-weight
GLM 5BudgetOPEN
Z.AIFeb 2026205K$0.60$1.92$0.1269.3 t/s1.4s
Z.AI · 205K context · Feb 2026 · cached $0.12 · open-weight · 69.3 t/s · 1.4s TTFT
MoonshotKimi K2.7 CodeBudgetOPEN
MoonshotJun 2026262K$0.66$3.40$0.18
Moonshot · 262K context · Jun 2026 · cached $0.18 · open-weight
GoogleMay 20261M$0.75$4.50$0.07
Google · 1M context · May 2026 · cached $0.07
GoogleJul 20261M$0.75$3.75$0.07
Google · 1M context · Jul 2026 · cached $0.07
OpenAIGPT-5.4 MiniBudget
OpenAIMar 2026400K$0.75$4.50$0.07180.7 t/s10.53s
OpenAI · 400K context · Mar 2026 · cached $0.07 · 180.7 t/s · 10.53s TTFT
QwenFeb 2026262K$0.78$3.90$0.08
Qwen · 262K context · Feb 2026 · cached $0.08
DeepSeekApr 20261M$0.87$1.74$0.07
DeepSeek · 1M context · Apr 2026 · cached $0.07 · open-weight
MoonshotKimi K2.6BudgetOPEN
MoonshotApr 2026262K$0.95$4.00$0.1644.3 t/s2.78s
Moonshot · 262K context · Apr 2026 · cached $0.16 · open-weight · 44.3 t/s · 2.78s TTFT
MoonshotJun 2026262K$0.95$4.00$0.19
Moonshot · 262K context · Jun 2026 · cached $0.19 · open-weight
AnthropicJun 20261M$1.00$5.00$0.10
Anthropic · 1M context · Jun 2026 · cached $0.10
GoogleFeb 20261M$1.00$6.00$0.10
Google · 1M context · Feb 2026 · cached $0.10
OpenAIJul 20261M$1.00$5.00$0.10
OpenAI · 1M context · Jul 2026 · cached $0.10
OpenAIJul 20261M$1.00$5.00$0.10
OpenAI · 1M context · Jul 2026 · cached $0.10
OpenAIJul 20261M$1.00$6.00$0.10
OpenAI · 1M context · Jul 2026 · cached $0.10
OpenAIJul 20261M$1.00$6.00$0.10
OpenAI · 1M context · Jul 2026 · cached $0.10
Grok Build 0.1Standard
xAIMay 2026256K$1.00$2.00$0.20
xAI · 256K context · May 2026 · cached $0.20
QwenApr 2026262K$1.03$6.16$0.10
Qwen · 262K context · Apr 2026 · cached $0.10
DeepSeekDeepSeek V4 Pro 0813StandardOPEN
DeepSeekAug 20261M$1.12$3.37$0.04
DeepSeek · 1M context · Aug 2026 · cached $0.04 · open-weight
GLM 5.2StandardOPEN
Z.AIJun 20261M$1.19$3.74$0.22
Z.AI · 1M context · Jun 2026 · cached $0.22 · open-weight
GLM 5 TurboStandard
Z.AIMar 2026203K$1.20$4.00$0.24
Z.AI · 203K context · Mar 2026 · cached $0.24
GLM 5V TurboStandard
Z.AIApr 2026203K$1.20$4.00$0.24
Z.AI · 203K context · Apr 2026 · cached $0.24
OpenAIGPT-5.4 (batch)Standard
OpenAIMar 20261M$1.25$7.50$0.13
OpenAI · 1M context · Mar 2026 · cached $0.13
Grok 4.20Standard
xAIMar 20262M$1.25$2.50$0.20
xAI · 2M context · Mar 2026 · cached $0.20
xAIMar 20262M$1.25$2.50$0.20
xAI · 2M context · Mar 2026 · cached $0.20
Grok 4.3Standard
xAIApr 20261M$1.25$2.50$0.20
xAI · 1M context · Apr 2026 · cached $0.20
GLM 5.1StandardOPEN
Z.AIApr 2026205K$1.26$3.96$0.23
Z.AI · 205K context · Apr 2026 · cached $0.23 · open-weight
GLM 5.3Standard
Z.AIAug 20261M$1.40$4.40$0.26
Z.AI · 1M context · Aug 2026 · cached $0.26
QwenQwen3.7 MaxStandard
QwenMay 20261M$1.48$4.42$0.29
Qwen · 1M context · May 2026 · cached $0.29
AnthropicFeb 20261M$1.50$7.50$0.15
Anthropic · 1M context · Feb 2026 · cached $0.15
GoogleGemini 3.5 FlashStandard
GoogleMay 20261M$1.50$9.00$0.15206.1 t/s15.35s
Google · 1M context · May 2026 · cached $0.15 · 206.1 t/s · 15.35s TTFT
MistralApr 2026262K$1.50$7.50$0.15
Mistral · 262K context · Apr 2026 · cached $0.15
OpenAIGPT-5.2-CodexStandard
OpenAIJan 2026400K$1.75$14.00$0.17
OpenAI · 400K context · Jan 2026 · cached $0.17
OpenAIGPT-5.3-CodexStandard
OpenAIFeb 2026400K$1.75$14.00$0.17
OpenAI · 400K context · Feb 2026 · cached $0.17
AnthropicClaude Sonnet 5Standard
AnthropicJun 20261M$2.00$10.00$0.20
Anthropic · 1M context · Jun 2026 · cached $0.20
GoogleFeb 20261M$2.00$12.00$0.20
Google · 1M context · Feb 2026 · cached $0.20
GoogleFeb 20261M$2.00$12.00$0.20
Google · 1M context · Feb 2026 · cached $0.20
OpenAIGPT-5.6 SolStandard
OpenAIJul 20261M$2.00$10.00$0.2074.6 t/s3.21s
OpenAI · 1M context · Jul 2026 · cached $0.20 · 74.6 t/s · 3.21s TTFT
OpenAIGPT-5.6 Sol ProStandard
OpenAIJul 20261M$2.00$10.00$0.20
OpenAI · 1M context · Jul 2026 · cached $0.20
OpenAIGPT-5.6 TerraStandard
OpenAIJul 20261M$2.00$12.00$0.20101.1 t/s1.8s
OpenAI · 1M context · Jul 2026 · cached $0.20 · 101.1 t/s · 1.8s TTFT
OpenAIJul 20261M$2.00$12.00$0.20
OpenAI · 1M context · Jul 2026 · cached $0.20
QwenQwen3.8 2.4T A95BStandardOPEN
QwenAug 20261M$2.00$6.00$0.25
Qwen · 1M context · Aug 2026 · cached $0.25 · open-weight
QwenQwen3.8 MaxStandard
QwenAug 20261M$2.00$6.00$0.25
Qwen · 1M context · Aug 2026 · cached $0.25
Grok 4.5Standard
xAIJul 2026500K$2.00$6.00$0.30
xAI · 500K context · Jul 2026 · cached $0.30
Grok 4.6Standard
xAIAug 2026500K$2.00$6.00$0.5059.4 t/s38.13s
xAI · 500K context · Aug 2026 · cached $0.50 · 59.4 t/s · 38.13s TTFT
AnthropicFeb 20261M$2.50$12.50$0.25
Anthropic · 1M context · Feb 2026 · cached $0.25
AnthropicApr 20261M$2.50$12.50$0.25
Anthropic · 1M context · Apr 2026 · cached $0.25
AnthropicMay 20261M$2.50$12.50$0.25
Anthropic · 1M context · May 2026 · cached $0.25
AnthropicJul 20261M$2.50$12.50$0.25
Anthropic · 1M context · Jul 2026 · cached $0.25
OpenAIGPT-5.4Standard
OpenAIMar 20261M$2.50$15.00$0.25
OpenAI · 1M context · Mar 2026 · cached $0.25
OpenAIGPT-5.5 (batch)Standard
OpenAIApr 20261M$2.50$15.00$0.25
OpenAI · 1M context · Apr 2026 · cached $0.25
AnthropicFeb 20261M$3.00$15.00$0.30
Anthropic · 1M context · Feb 2026 · cached $0.30
MoonshotKimi K3StandardOPEN
MoonshotJul 20261M$3.00$15.00$0.3037.5 t/s4.05s
Moonshot · 1M context · Jul 2026 · cached $0.30 · open-weight · 37.5 t/s · 4.05s TTFT
AnthropicJun 20261M$5.00$25.00$0.50
Anthropic · 1M context · Jun 2026 · cached $0.50
AnthropicClaude Opus 4.6Premium
AnthropicFeb 20261M$5.00$25.00$0.50
Anthropic · 1M context · Feb 2026 · cached $0.50
AnthropicClaude Opus 4.7Premium
AnthropicApr 20261M$5.00$25.00$0.50
Anthropic · 1M context · Apr 2026 · cached $0.50
AnthropicClaude Opus 4.8Premium
AnthropicMay 20261M$5.00$25.00$0.50
Anthropic · 1M context · May 2026 · cached $0.50
AnthropicClaude Opus 5Premium
AnthropicJul 20261M$5.00$25.00$0.50
Anthropic · 1M context · Jul 2026 · cached $0.50
OpenAIGPT Chat LatestPremium
OpenAIMay 2026400K$5.00$30.00$0.50
OpenAI · 400K context · May 2026 · cached $0.50
OpenAIGPT-5.5Premium
OpenAIApr 20261M$5.00$30.00$0.5088 t/s61.39s
OpenAI · 1M context · Apr 2026 · cached $0.50 · 88 t/s · 61.39s TTFT
AnthropicClaude Fable 5Premium
AnthropicJun 20261M$10.00$50.00$1.00
Anthropic · 1M context · Jun 2026 · cached $1.00
AnthropicMay 20261M$10.00$50.00$1.00
Anthropic · 1M context · May 2026 · cached $1.00
AnthropicJul 20261M$10.00$50.00$1.00
Anthropic · 1M context · Jul 2026 · cached $1.00
OpenAIMar 20261M$15.00$90.00$1.50
OpenAI · 1M context · Mar 2026 · cached $1.50
OpenAIApr 20261M$15.00$90.00$1.50
OpenAI · 1M context · Apr 2026 · cached $1.50
AnthropicMay 20261M$30.00$150.00$3.00
Anthropic · 1M context · May 2026 · cached $3.00
OpenAIGPT-5.4 ProPremium
OpenAIMar 20261M$30.00$180.00$3.00
OpenAI · 1M context · Mar 2026 · cached $3.00
OpenAIGPT-5.5 ProPremium
OpenAIApr 20261M$30.00$180.00$3.00
OpenAI · 1M context · Apr 2026 · cached $3.00
108 / 108 modelsCached-input rate shown per model.

Figures are refreshed daily from live provider pricing (via OpenRouter), for planning only. Always confirm current rates on the provider's own pricing page before committing spend. Speed (tokens/sec, time-to-first-token) data from Artificial Analysis. Reusing a system prompt or document across requests? See which model discounts caching the most on the cache savings calculator.

Understanding LLM API pricing

Large language model providers bill by the token — a chunk of text roughly ¾ of a word in English (about 4 characters). Every request has two metered parts: the input (your prompt, system instructions, and any context you send) and the output (the text the model generates back). Providers publish separate per-token rates for each, and those rates differ by orders of magnitude across models — which is exactly why an apples-to-apples table is useful before you commit to one.

Why we normalize to “per 1M tokens”

Providers quote prices in different units — per 1,000 tokens, per million tokens, sometimes per character. TokenCost converts every model to a single unit, US dollars per one million tokens, so you can line up a budget model against a frontier model and compare the numbers directly. All figures on this page use that unit.

Input, output, and cached-input rates

  • Input (prompt) price — what you pay per 1M tokens you send to the model. Long system prompts, retrieved documents, and chat history all count as input.
  • Output (completion) price — what you pay per 1M tokens the model returns. This is almost always the more expensive of the two, often 3–8× the input rate, because generation is the compute-heavy step. If your app produces long answers, output cost usually dominates your bill.
  • Cached input — many providers now offer prompt caching: if you resend an identical prefix (a fixed system prompt, a large document) within a short window, the repeated tokens are billed at a steep discount. The table shows each model’s own cached-input rate rather than a flat estimate, because the discount varies widely by provider.

Context window and tiers

The context window is the maximum number of tokens (input + output) a model can consider at once — larger windows let you send more documents or longer conversations, but sending more tokens costs more. We also tag each model with a rough tier — Budget, Standard, or Premium — as a quick signal of where it sits on the price/capability spectrum, so you can filter to the class you actually need.

How to use this page

Filter by provider, switch the sort (cheapest input, cheapest output, largest context, newest release), or narrow by release year. On a phone the table condenses to the essentials — tap any row to expand the full detail. Once you’ve found candidates, project a real monthly bill on the cost calculator, price an exact prompt with the token counter, find the lowest-cost options on cheapest models, put finalists head-to-head on the compare page, check capability against price on the benchmarks, or see every price cut and increase we've logged on price trends. For where the numbers come from and how often they refresh, see our methodology.

Frequently asked questions

What is a token?
A token is a chunk of text a model reads or writes — in English, roughly ¾ of a word or about 4 characters. Both the text you send (input) and the text the model returns (output) are counted in tokens, and you are billed per token for each.
Why are output tokens more expensive than input tokens?
Generating text is the compute-intensive step for a language model, so providers price output (completion) tokens higher than input (prompt) tokens — commonly 3 to 8 times higher. If your application returns long responses, output cost usually dominates the bill.
What is cached input, or prompt caching?
Prompt caching lets you reuse an identical prompt prefix — such as a fixed system prompt or a large reference document — at a large discount when it is resent within a provider's caching window. TokenCost shows each model's own cached-input rate rather than a flat assumption.
How often is the pricing updated?
Prices are refreshed automatically every day from live provider data via OpenRouter, then rebuilt into the site. The “updated” date in the index summary shows the latest refresh.
Which providers and models are included?
The table tracks current models from the major API providers — including OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, Moonshot and Qwen. Models are added automatically as providers ship them.
Are these the official prices? Can I rely on them for billing?
TokenCost is a planning aid. Figures are aggregated from live data and refreshed daily, but you should always confirm the current rate on the provider’s own pricing page before committing spend. We are not affiliated with any provider.