Last verified
50% OFF TO SEP 9NATIVE MULTIMODAL320B / 18B ACTIVE1M CONTEXTTHINKING ALWAYS ON

GLM-5.3-Flash API Pricing

GLM-5.3-Flash is the first native multimodal model in the GLM-5 line, and it is priced like a small model without being one: 320B parameters with 18B active, listed at $0.15/M input and $0.50/M output — roughly a ninth of what GLM-5.3 costs. Z.ai is running a 50% discount on that list until September 9, 2026, so a call made today bills at $0.075/M input and $0.25/M output. The figures on this page are list price.

Input - per 1M tokens
$0.15/M
Base token price - $0.075 until Sep 9 list price
Output - per 1M tokens
$0.50/M
Output tokens - $0.25 until Sep 9 list price
Cached input - per 1M tokens
$0.03/M
Cache hit, 5x cheaper than base -80%
Effective - agentic blend
$0.0875/M
92/8 split - 82% cache hit rate
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with Z.ai's International list rates. The promotional half-price runs to 24:00 on September 9, 2026 (UTC+8) — halve every figure below to see today's bill. Thinking cannot be switched off on this model, so reasoning tokens land in the output count on every call.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (zai-org/GLM-5, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 469
Words 82
Tokens (estimated) 89 tokens
Cost as input · uncached $0.00001 USD
Cost as output · uncached $0.00004 USD
Cost as cached input $0.000003 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
GLM-5.3-Flash Current $0.15 cache $0.03 $0.50 $0.0875 list price 1M Multimodal frontier work at a light-tier rate
GLM-4.7-Flash Free cache Free Free Free free tier 200K Free, and a generation behind
GLM-4.7-FlashX $0.07 cache $0.01 $0.40 $0.0511 cheaper 200K Cheapest paid GLM, no multimodal
GLM-4.7 $0.60 cache $0.11 $2.20 $0.358 pricier 200K Previous mid tier, fifth the context
GLM-5.3 $1.40 cache $0.26 $4.40 $0.78 pricier 1M The full-size flagship, text only
Qwen3.8-Flash $0.15 $0.47 $0.176 pricier 1M Same list input, no published cache rate
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 pricier 1M The Western flash tier, five times the rate
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

GLM-5.3-Flash pricing questions, with the list price kept apart from the discount running on top of it.

Q · 01 How much does GLM-5.3-Flash cost? +
List price is $0.15/M input, $0.03/M on a cache hit and $0.50/M output. Until 24:00 on September 9, 2026 (UTC+8) Z.ai is charging half of each: $0.075, $0.015 and $0.25. This catalogue stores list prices, so every figure on this page is the undiscounted rate.
Q · 02 Why store the list price rather than what I pay today? +
Because a promotion that ends quietly would otherwise turn every stored figure false overnight, and because comparisons across 300 rows only hold if they are all on the same basis. Chinese vendors run these discounts constantly and often without an end date; recording the list keeps the catalogue consistent and errs toward over-estimating your bill rather than under-estimating it. The discounted rate is recorded alongside, and this page prints both.
Q · 03 How does it compare with GLM-5.3? +
GLM-5.3 lists at $1.4/M input and $4.4/M output — 9.3x and 8.8x this model's list rate. On the agentic blend the gap is 8.9x. The Flash is also the more capable model in one respect the flagship cannot match: it is natively multimodal, where GLM-5.3 is text in and text out.
Q · 04 Can thinking be turned off? +
No. Z.ai documents that thinking.type only supports enabled on this model, so reasoning tokens are billed as output on every call. Budget the output side more generously than a non-reasoning model of the same price would need — the vendor recommends reasoning_effort: max and setting thinking.clear_thinking to false.
Q · 05 What does it accept as input? +
Images, video and files, natively — Z.ai calls it "the first native multimodal model in the GLM-5 series" and the visual capability is built into the coding loop rather than bolted on, so the model can look at rendered output and iterate. Text parameters match GLM-5.3, including the 1M-token context window.
Q · 06 What about the free cached-input storage? +
Z.ai marks cached input storage as limited-time free across the whole GLM line. That is a separate charge from the token rates — a fee for holding the cache, not for reading it — and this catalogue does not model it. The $0.03/M cache-hit rate above is unaffected by it.