Tell it what you’re building and what you can’t compromise on, and get a ranked shortlist from 58 models — with the reasoning spelled out and a monthly cost estimate on every pick.
| # | MODEL | TIER | BLENDED $/1M | CONTEXT |
|---|---|---|---|---|
| 4 | Grok 4.6 xAI | FRONTIER | $3.00 | 500K |
| 5 | Grok 4.5 xAI | FRONTIER | $3.00 | 500K |
| 6 | Qwen3.8-Max Alibaba | FRONTIER | $3.00 | 1M |
| 7 | Mistral Medium 3.5 Mistral | FRONTIER | $3.00 | 262K |
| 8 | Claude Sonnet 5 Anthropic | MID | $4.00 | 1M |
| 9 | GPT-5.6 Terra OpenAI | MID | $4.50 | 1.05M |
| 10 | Gemini 3.1 Pro Google | FRONTIER | $4.50 | 1.05M |
| 11 | Kimi K3 Moonshot AI | FRONTIER | $6.00 | 1.05M |
Blended assumes a 3:1 input:output ratio. Prices verified August 31, 2026 — see the full pricing table.
Must-have capabilities and the context floor are non-negotiable — a model either qualifies or it doesn’t. The survivors are ranked by capability tier and blended price, weighted by your priority. “Balanced” rewards cheap frontier models hardest, which is why picks like DeepSeek and GLM show up more often than their marketing budgets would suggest.
Use-case presets also carry a short, curated list of models with a proven track record for that job. When one of those influenced a ranking, the first “why” bullet says so — no silent thumbs on the scale.
Treat the shortlist as a starting bench, not a verdict: run your own evals on the top two or three before committing.
Your must-have capabilities and minimum context act as hard filters. What remains is scored by capability tier and blended price per million tokens, weighted by your priority — lowest cost sorts purely by price, max capability by tier first, balanced by capability per dollar. Some use cases carry a curated boost for models with a proven track record there; when a boost influenced a pick, the “why” bullets say so.
Frontier is the top capability class — the models labs put their best work into (Claude Opus, GPT-5.6 Sol, Gemini Pro, DeepSeek V4 Pro). Mid is the workhorse tier: strong quality at a fraction of frontier prices. Small is built for speed and volume — classification, extraction, and high-throughput tasks where per-token cost dominates.
For agentic coding — multi-file edits, running tools, long sessions — frontier reasoning models like Claude Opus 5 and GPT-5.6 Sol lead. On a budget, coding-tuned open models such as Kimi K2.7 Code and xAI’s Grok Build 0.1 punch far above their price. Pick the Coding preset above and the ranking adjusts to exactly this.
The small tiers of the big labs (GPT-5.6 Luna, Claude Haiku 4.5, Gemini Flash-Lite) are shockingly capable for cents per million tokens, and open-weight models served via API — GLM-5.3-Flash, DeepSeek V4 Flash — often undercut them further. Set the priority to “Lowest cost” and compare the survivors.
Prices come from the same dataset as our pricing table, checked against each provider’s official pricing page — the stamp at the top shows the verification date. Spot something stale? Email support@reqkey.com.
ReqKey gives every API key a credit balance — validate, deduct, enforce limits, and watch per-consumer AI spend in real time.