FREE TOOL · PRICES VERIFIED AUGUST 31, 2026

GLM API Pricing

A frontier open model at $1.40 per million input tokens, and a multimodal Flash tier priced close to free. Real rates for every GLM model, a workload calculator, and the head-to-head comparisons.

MODELS
3
INPUT FROM
$0.075/1M
OUTPUT UP TO
$4.40/1M
MAX CONTEXT
1.31M
ESTIMATE YOUR WORKLOAD
MODELINPUT $/1MCACHED INOUTPUT $/1MBLENDEDCONTEXTYOUR COST
GLM-5.3
glm-5.3 · flagship
$1.40$0.26$4.40$2.151.31M$3.60
GLM-5.3-Flash
glm-5.3-flash
$0.075$0.015$0.25$0.1191.31M$0.20
GLM-4.7
glm-4.7 · prev gen
$0.60$0.11$2.20$1.00205K$1.70

USD per 1M tokens from Z.ai’s official pricing page as of August 31, 2026. Blended assumes a 3:1 input:output ratio. Get keys at the Z.ai console.

HOW Z.AI PRICING WORKS

A frontier model and a near-free one, with a modality twist.

GLM-5.3 at $1.40 input / $4.40 output per 1M tokens is among the cheapest frontier-class stickers anywhere, with a 1.31M-token context and open weights. The twist to know: GLM-5.3 itself is text-only, while GLM-5.3-Flash is the multimodal one — it takes images and video. Pick by modality, not just by budget.

And Flash is the headline bargain of this whole site: $0.075 input / $0.25 output per 1M tokens for a multimodal model with the same 1.31M window. For bulk classification, extraction, and vision workloads, nothing at this capability level comes close on price.

GLM-4.7 stays available at $0.60 / $2.20 for workloads pinned to the previous generation. Cache reads bill at a fraction of fresh input across the line ($0.26 on GLM-5.3), and the open weights mean self-hosting and third-party hosts are always an option.

How much does the GLM API cost?

GLM-5.3 costs $1.40 input / $4.40 output per 1M tokens, GLM-5.3-Flash $0.075 / $0.25, and the previous-generation GLM-4.7 $0.60 / $2.20. Cache reads cost a fraction of fresh input on all three.

What is the difference between GLM-5.3 and GLM-5.3-Flash?

Capability class and modality — and the modality goes the way you might not expect. GLM-5.3 is the frontier reasoning model but takes text only; GLM-5.3-Flash is smaller and 18× cheaper, yet accepts images and video. Vision workloads go to Flash; hard text reasoning goes to 5.3.

Is GLM cheaper than DeepSeek?

They are close at the top — GLM-5.3 blends to about $2.15/1M versus $1.98 for DeepSeek V4 Pro (before DeepSeek’s off-peak discount). At the fast tier GLM wins on price and modality: Flash at $0.12/1M blended versus $0.66 for V4 Flash, with vision included. The matchup table below has both.

Can I self-host GLM models?

Yes — the GLM line ships open weights, so you can run it on your own GPUs or through third-party inference providers. The prices here are Z.ai’s first-party API rates.

Where do I get a GLM API key?

From Z.ai’s developer platform (docs.z.ai) — pay-as-you-go per token. GLM models are also widely available on OpenRouter under the z-ai prefix.

HOW Z.AI STACKS UP

GLM vs. its closest rivals, tier by tier

BLENDED $/1M · 3:1 IN:OUT
OPEN FRONTIER
GLM-5.3
$2.15/1M
DeepSeek V4 Pro
$1.98/1M
Kimi K3
$6.00/1M
COMPARE
VS CLOSED MID-TIER
GLM-5.3
$2.15/1M
Claude Sonnet 5
$4.00/1M
GPT-5.6 Terra
$4.50/1M
COMPARE
BUDGET MULTIMODAL
GLM-5.3-Flash
$0.119/1M
DeepSeek V4 Flash
$0.66/1M
Gemini 3.5 Flash-Lite
$0.85/1M
COMPARE

Also see: OpenAI API pricing · Claude API pricing · Gemini API pricing · Grok API pricing · DeepSeek API pricing · Kimi API pricing · MiniMax API pricing · MiMo API pricing · all 58 models

BUILDING ON THE Z.AI API?

Your users will spend these tokens. Meter them per API key.

ReqKey gives every key a credit balance — validate, deduct, enforce limits, and watch per-consumer GLM spend in real time.