Every current Gemini model with its real per-million-token price — plus the free tier, the long-context surcharge, a calculator for your own workload, and a straight comparison against GPT and Claude.
| MODEL | INPUT $/1M | CACHED IN | OUTPUT $/1M | BLENDED | CONTEXT | YOUR COST |
|---|---|---|---|---|---|---|
Gemini 3.1 Pro gemini-3.1-pro-preview · >200K: $4/$18 | $2.00 | $0.20 | $12.00 | $4.50 | 1.05M | $8.00 |
Gemini 3.7 Flash gemini-3.7-flash · promo price | $0.75 | $0.075 | $3.75 | $1.50 | 1.05M | $2.63 |
Gemini 3.5 Flash gemini-3.5-flash · prior gen | $1.50 | $0.15 | $9.00 | $3.38 | 1.05M | $6.00 |
Gemini 3.5 Flash-Lite gemini-3.5-flash-lite · cheapest tier | $0.30 | $0.03 | $2.50 | $0.85 | 1.05M | $1.55 |
USD per 1M tokens from Google’s official pricing page as of August 31, 2026. Blended assumes a 3:1 input:output ratio. Get keys at the Google console.
Gemini is the only one of the big three with a genuinely free API tier: AI Studio keys run at no cost within rate limits, which is why so many prototypes start here. Paid rates are aggressive too — Gemini 3.7 Flash at $0.75 input / $3.75 output per 1M tokens is the cheapest current-generation multimodal workhorse on the market, and every current Gemini model takes text, images, audio, and video in a 1M-token context.
The catch to budget for: long-context surcharges. Prompts beyond 200K tokens bill at a higher rate on Gemini 3.1 Pro — $4 input / $18 output instead of $2 / $12. If you routinely stuff half the context window, your effective rate is closer to double the sticker price.
Context caching bills cache reads at about a tenth of the fresh input rate, plus a storage fee of $0.50 per million tokens per hour for explicit caches (implicit caching kicks in automatically on prompts over ~4K tokens). The Batch API halves prices for asynchronous jobs — worth it for bulk multimodal processing, where Gemini is usually the cheapest option anyway.
Gemini bills per token with separate input and output rates per million tokens: Gemini 3.1 Pro at $2 / $12 (rising to $4 / $18 beyond 200K prompt tokens), 3.7 Flash at $0.75 / $3.75, and Flash-Lite at $0.30 / $2.50. All current models handle text, images, audio, and video.
Yes, within limits: AI Studio API keys have a free tier with rate limits that is enough for prototypes and small tools. Production traffic needs the paid tier, billed per token at the rates in the table above.
On Gemini 3.1 Pro, any request whose prompt exceeds 200K tokens bills the whole request at the higher tier — $4 input / $18 output instead of $2 / $12. Flash models keep flat pricing. If you regularly send huge contexts, price your workload at the surcharge rate.
Usually, tier for tier: Gemini 3.1 Pro ($4.50/1M blended) undercuts GPT-5.6 Sol ($8) and Claude Opus 5 ($10), and 3.7 Flash ($1.50) undercuts both mid-tier rivals — as long as you stay under the 200K long-context threshold. The matchup table below shows all three side by side.
Start with 3.7 Flash — current generation, fully multimodal, and a fraction of Pro pricing. Move up to 3.1 Pro when answer quality on hard reasoning is worth 3× the cost, or down to Flash-Lite for bulk classification at $0.30 / $2.50.
Also see: OpenAI API pricing · Claude API pricing · Grok API pricing · DeepSeek API pricing · Kimi API pricing · GLM API pricing · MiniMax API pricing · MiMo API pricing · all 58 models
ReqKey gives every key a credit balance — validate, deduct, enforce limits, and watch per-consumer Gemini spend in real time.