A frontier open model at $1.40 per million input tokens, and a multimodal Flash tier priced close to free. Real rates for every GLM model, a workload calculator, and the head-to-head comparisons.
| MODEL | INPUT $/1M | CACHED IN | OUTPUT $/1M | BLENDED | CONTEXT | YOUR COST |
|---|---|---|---|---|---|---|
GLM-5.3 glm-5.3 · flagship | $1.40 | $0.26 | $4.40 | $2.15 | 1.31M | $3.60 |
GLM-5.3-Flash glm-5.3-flash | $0.075 | $0.015 | $0.25 | $0.119 | 1.31M | $0.20 |
GLM-4.7 glm-4.7 · prev gen | $0.60 | $0.11 | $2.20 | $1.00 | 205K | $1.70 |
USD per 1M tokens from Z.ai’s official pricing page as of August 31, 2026. Blended assumes a 3:1 input:output ratio. Get keys at the Z.ai console.
GLM-5.3 at $1.40 input / $4.40 output per 1M tokens is among the cheapest frontier-class stickers anywhere, with a 1.31M-token context and open weights. The twist to know: GLM-5.3 itself is text-only, while GLM-5.3-Flash is the multimodal one — it takes images and video. Pick by modality, not just by budget.
And Flash is the headline bargain of this whole site: $0.075 input / $0.25 output per 1M tokens for a multimodal model with the same 1.31M window. For bulk classification, extraction, and vision workloads, nothing at this capability level comes close on price.
GLM-4.7 stays available at $0.60 / $2.20 for workloads pinned to the previous generation. Cache reads bill at a fraction of fresh input across the line ($0.26 on GLM-5.3), and the open weights mean self-hosting and third-party hosts are always an option.
GLM-5.3 costs $1.40 input / $4.40 output per 1M tokens, GLM-5.3-Flash $0.075 / $0.25, and the previous-generation GLM-4.7 $0.60 / $2.20. Cache reads cost a fraction of fresh input on all three.
Capability class and modality — and the modality goes the way you might not expect. GLM-5.3 is the frontier reasoning model but takes text only; GLM-5.3-Flash is smaller and 18× cheaper, yet accepts images and video. Vision workloads go to Flash; hard text reasoning goes to 5.3.
They are close at the top — GLM-5.3 blends to about $2.15/1M versus $1.98 for DeepSeek V4 Pro (before DeepSeek’s off-peak discount). At the fast tier GLM wins on price and modality: Flash at $0.12/1M blended versus $0.66 for V4 Flash, with vision included. The matchup table below has both.
Yes — the GLM line ships open weights, so you can run it on your own GPUs or through third-party inference providers. The prices here are Z.ai’s first-party API rates.
From Z.ai’s developer platform (docs.z.ai) — pay-as-you-go per token. GLM models are also widely available on OpenRouter under the z-ai prefix.
Also see: OpenAI API pricing · Claude API pricing · Gemini API pricing · Grok API pricing · DeepSeek API pricing · Kimi API pricing · MiniMax API pricing · MiMo API pricing · all 58 models
ReqKey gives every key a credit balance — validate, deduct, enforce limits, and watch per-consumer GLM spend in real time.