Frontier-class open-weight models at commodity prices — plus an off-peak window that halves them again. Real per-million-token rates, a calculator for your workload, and the comparison against closed rivals.
| MODEL | INPUT $/1M | CACHED IN | OUTPUT $/1M | BLENDED | CONTEXT | YOUR COST |
|---|---|---|---|---|---|---|
DeepSeek V4 Pro deepseek-v4-pro-0813 · off-peak −50% | $1.32 | $0.044 | $3.96 | $1.98 | 1M | $3.30 |
DeepSeek V4 Flash deepseek-v4-flash-0731 · off-peak −50% | $0.44 | $0.014 | $1.32 | $0.66 | 1M | $1.10 |
USD per 1M tokens from DeepSeek’s official pricing page as of August 31, 2026. Blended assumes a 3:1 input:output ratio. Get keys at the DeepSeek console.
DeepSeek keeps the menu short: V4 Pro, the frontier reasoning model, at $1.32 input / $3.96 output per 1M tokens with a 1M context, and V4 Flash at $0.44 / $1.32 for volume work. Both are text-only (pair them with a vision model if you need images) and both ship open weights, so you can also self-host or buy them from any inference provider.
The signature mechanic is the off-peak discount: between 16:30 and 00:30 UTC, everything bills at half price. Schedule batch jobs, evaluations, and backfills into that window and you get a batch-API-grade discount without a batch API — V4 Pro effectively drops to $0.66 / $1.98.
Cache hits are the other outlier: repeated prompt prefixes bill at $0.044 per 1M tokens on V4 Pro — about a thirtieth of the fresh rate, the deepest cache discount on this site. Agents with long stable system prompts run absurdly cheap here.
DeepSeek V4 Pro costs $1.32 input / $3.96 output per 1M tokens, and V4 Flash $0.44 / $1.32 — with everything at half price during the off-peak window (16:30–00:30 UTC) and cache hits at roughly a thirtieth of the fresh input rate.
Requests processed between 16:30 and 00:30 UTC bill at 50% of the standard rate, input and output alike. There is nothing to configure — the discount applies automatically based on when the request lands. Time-shift bulk workloads into the window and your effective price halves.
V4 Pro sits in the frontier class on reasoning and coding benchmarks while costing a fraction of closed rivals — around $1.98/1M blended versus $8 for GPT-5.6 Sol and $10 for Claude Opus 5. The trade-offs: text-only input, and you are buying from a Chinese lab, which matters for some compliance postures.
Yes — the V4 models ship open weights, so you can run them on your own GPUs or through any inference provider. The prices on this page are DeepSeek’s first-party API rates, which are usually hard to beat at small scale once you count GPU costs.
The V4 API models priced here are text-only. If your workload needs image, audio, or video input at a similar price point, look at GLM-5.3-Flash, MiMo-V2.5, or Gemini Flash-Lite — the matchup table below includes the closest alternatives.
Also see: OpenAI API pricing · Claude API pricing · Gemini API pricing · Grok API pricing · Kimi API pricing · GLM API pricing · MiniMax API pricing · MiMo API pricing · all 58 models
ReqKey gives every key a credit balance — validate, deduct, enforce limits, and watch per-consumer DeepSeek spend in real time.