GLM-5.3-Flash API Pricing
GLM-5.3-Flash is the first native multimodal model in the GLM-5 line, and it is priced like a small model without being one: 320B parameters with 18B active, listed at $0.15/M input and $0.50/M output — roughly a ninth of what GLM-5.3 costs. Z.ai is running a 50% discount on that list until September 9, 2026, so a call made today bills at $0.075/M input and $0.25/M output. The figures on this page are list price.
Run the numbers.
Calculator pre-loaded with Z.ai's International list rates. The promotional half-price runs to 24:00 on September 9, 2026 (UTC+8) — halve every figure below to see today's bill. Thinking cannot be switched off on this model, so reasoning tokens land in the output count on every call.
Real-world presets.
Repo-wide feature build
Reading 150-page contracts
Support agent ticket triage
Research planning turn
Paste text. See tokens. See cost.
Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (zai-org/GLM-5, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| GLM-5.3-Flash Current | $0.15 cache $0.03 | $0.50 | $0.0875 list price | 1M | Multimodal frontier work at a light-tier rate |
| GLM-4.7-Flash | Free cache Free | Free | Free free tier | 200K | Free, and a generation behind |
| GLM-4.7-FlashX | $0.07 cache $0.01 | $0.40 | $0.0511 cheaper | 200K | Cheapest paid GLM, no multimodal |
| GLM-4.7 | $0.60 cache $0.11 | $2.20 | $0.358 pricier | 200K | Previous mid tier, fifth the context |
| GLM-5.3 | $1.40 cache $0.26 | $4.40 | $0.78 pricier | 1M | The full-size flagship, text only |
| Qwen3.8-Flash | $0.15 | $0.47 | $0.176 pricier | 1M | Same list input, no published cache rate |
| Gemini 3.7 Flash | $0.75 cache $0.075 | $3.75 | $0.481 pricier | 1M | The Western flash tier, five times the rate |
Frequently asked.
GLM-5.3-Flash pricing questions, with the list price kept apart from the discount running on top of it.
Q · 01 How much does GLM-5.3-Flash cost? +
$0.15/M input, $0.03/M on a cache hit and $0.50/M output. Until 24:00 on September 9, 2026 (UTC+8) Z.ai is charging half of each: $0.075, $0.015 and $0.25. This catalogue stores list prices, so every figure on this page is the undiscounted rate.Q · 02 Why store the list price rather than what I pay today? +
Q · 03 How does it compare with GLM-5.3? +
$1.4/M input and $4.4/M output — 9.3x and 8.8x this model's list rate. On the agentic blend the gap is 8.9x. The Flash is also the more capable model in one respect the flagship cannot match: it is natively multimodal, where GLM-5.3 is text in and text out.Q · 04 Can thinking be turned off? +
thinking.type only supports enabled on this model, so reasoning tokens are billed as output on every call. Budget the output side more generously than a non-reasoning model of the same price would need — the vendor recommends reasoning_effort: max and setting thinking.clear_thinking to false.Q · 05 What does it accept as input? +
Q · 06 What about the free cached-input storage? +
$0.03/M cache-hit rate above is unaffected by it.