FREE TOOL · PRICES VERIFIED AUGUST 30, 2026

LLM API Pricing Table

What every major model actually costs per million tokens — GPT, Claude, Gemini, Grok, DeepSeek, Kimi, GLM and more. Sort any column, plug in your own workload, and compare models side by side.

ESTIMATE YOUR WORKLOAD
= 1,000 in + 500 out tokens × 1,000 requests. The your cost column updates live.
FILTER BY CAPABILITY
58 MODELS · 21 PROVIDERS · TICK UP TO 3 TO COMPARE
LINKS
GPT-5.6 Sol
OpenAI · flagship
$4.00$0.40$20.001.05M$14.00CONSOLEOPENROUTER
GPT-5.6 Terra
OpenAI
$2.00$0.20$12.001.05M$8.00CONSOLEOPENROUTER
GPT-5.5
OpenAI · prev flagship
$5.00$0.50$30.001.05M$20.00CONSOLEOPENROUTER
GPT-5.6 Luna
OpenAI · fast tier
$0.20$0.02$1.201.05M$0.80CONSOLEOPENROUTER
GPT-5.4 Mini
OpenAI
$0.75$0.075$4.50400K$3.00CONSOLEOPENROUTER
GPT-5.4 Nano
OpenAI
$0.20$1.25400K$0.82CONSOLEOPENROUTER
Claude Fable 5
Anthropic · top capability
$10.00$1.00$50.001M$35.00CONSOLEOPENROUTER
Claude Opus 5
Anthropic · flagship
$5.00$0.50$25.001M$17.50CONSOLEOPENROUTER
Claude Opus 4.8
Anthropic · prev flagship
$5.00$0.50$25.001M$17.50CONSOLEOPENROUTER
Claude Sonnet 5
Anthropic
$2.00$0.20$10.001M$7.00CONSOLEOPENROUTER
Claude Haiku 4.5
Anthropic · fast tier
$1.00$0.10$5.00200K$3.50CONSOLEOPENROUTER
Gemini 3.1 Pro
Google · >200K: $4/$18
$2.00$0.20$12.001.05M$8.00CONSOLEOPENROUTER
112 OF 58

All prices in USD per 1M tokens, from official provider pricing pages as of August 30, 2026. Tiered long-context surcharges noted per model. Spot an outdated price? Tell us.

HOW TOKEN PRICING WORKS

You pay per token, in and out — and output is the expensive direction.

Every LLM API bill is the same formula: input tokens × input rate + output tokens × output rate. Rates are quoted per million tokens, and output typically costs 3–6× more than input. Long system prompts, retries, and verbose responses are where budgets quietly die.

Cached input is the biggest lever most teams ignore: resend the same prompt prefix and providers charge roughly a tenth of the fresh-input rate. If you run agents or RAG with a long, stable system prompt, cache pricing — not model choice — often decides your bill.

Prices moved fast this year, almost always downward. We verify every number on this page against the provider’s official pricing docs and stamp the date up top.

How is LLM API usage priced?

Every major provider bills per token, with separate rates for input (your prompt) and output (the model’s reply). Prices are quoted per million tokens. Output tokens usually cost 3–6× more than input tokens, so response length drives most of the bill for chat-style workloads.

What exactly is a token?

A token is a chunk of text — roughly 4 characters or ¾ of an English word. “How much does the Claude API cost?” is about 9 tokens. A million tokens is roughly 750,000 words, or about ten novels.

What is cached-input pricing?

If you resend the same prompt prefix (a system prompt, few-shot examples, long documents), providers can serve it from cache at a steep discount — typically 10× cheaper than fresh input tokens. Agents and RAG apps with long stable prefixes save the most.

Which LLM API is the cheapest in 2026?

Depends on the tier you need. Among frontier flagships, mid-tier models like Claude Sonnet 5 and GPT-5.6 Terra cost a fraction of the top models, and small-tier models (GPT-5.6 Luna, Gemini Flash-Lite, Haiku 4.5) are 10–50× cheaper still. Open-weight models served via API are often the lowest ¢/token — sort the table by output price and compare against your quality bar.

How often is this table updated?

Prices are verified against each provider’s official pricing page — last check August 30, 2026. Vendors change prices often (usually downward, sometimes with promo windows). Spot something stale? Email support@reqkey.com and we’ll fix it within a day.

WHY WE BUILT THIS

Knowing the price is step one. Metering it per user is the hard part.

ReqKey gives every API key a credit balance — validate a key, deduct credits, enforce limits, and see spend per consumer in real time. One call from your own middleware.