Last verified
SAME CARD AS GLM-5.21M CONTEXT128K MAX OUTPUTTEXT ONLYPOST-TRAINING GAINS

GLM-5.3 API Pricing

GLM-5.3 is priced. Z.ai's International rate card lists $1.4/M input, $4.4/M output and $0.26/M cached input — the identical card GLM-5.2 and GLM-5.1 carry. That consistency is not a coincidence: Z.ai states GLM-5.3 uses the same base model as GLM-5.2, with every gain coming from post-training. A new generation at no extra cost per token. Until August 16, 2026 this page showed TBA — the API was listed as coming soon and access ran through the GLM Coding Plan subscription; the rate card now lists it.

Input - per 1M tokens
$1.40/M
Source Z.AI flat
Output - per 1M tokens
$4.40/M
Context 1M flat
Cached input - per 1M tokens
$0.26/M
Storage limited-time free -81%
Effective - agentic blend
$0.78/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with the live GLM-5.3 rates. Because the card matches GLM-5.2 exactly, every figure here doubles as a GLM-5.2 estimate — the choice between them is about capability, not cost.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (zai-org/GLM-5, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 420
Words 68
Tokens (estimated) 80 tokens
Cost as input · uncached $0.00011 USD
Cost as output · uncached $0.00035 USD
Cost as cached input $0.00002 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
GLM-5.3 Current $1.40 cache $0.26 $4.40 $0.78 current page 1M Newest GLM flagship
GLM-5.2 $1.40 cache $0.26 $4.40 $0.78 same base model 1M The generation underneath 5.3
GLM-5.1 $1.40 cache $0.26 $4.40 $0.78 same rate 200K Older GLM-5, narrower window
GLM-5-Turbo $1.20 cache $0.24 $4.00 $0.70 cheaper 200K Latency-sensitive GLM-5 work
GLM-5 $1.00 cache $0.20 $3.20 $0.572 cheaper 200K Cheapest of the GLM-5 line
GLM-4.7-FlashX $0.07 cache $0.01 $0.40 $0.0511 cheaper 200K High-throughput classification
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

GLM-5.3 pricing questions, including what changed when the API opened.

Q · 01 How much does GLM-5.3 cost? +
$1.4/M input, $4.4/M output and $0.26/M on a cache hit, on Z.ai's International rate card. Under AI//COST's 92/8 agentic blend with 82% cache hits, the effective planning figure is $0.78/M.
Q · 02 Is it more expensive than GLM-5.2? +
No — the two cards are identical, down to the cached rate. Z.ai states GLM-5.3 uses the same base model as GLM-5.2 and that the improvements come from post-training, so there is no price step between the generations. On cost alone there is no reason to stay on 5.2.
Q · 03 What changed since this page said TBA? +
The API opened. On August 16, 2026 GLM-5.3 had a model page marked New and shipped to GLM Coding Plan subscribers, while its own docs said "The GLM-5.3 API is coming soon" and neither the International nor the China rate card listed it — so this page showed TBA rather than guessing. As of August 23 the International card lists it.
Q · 04 How much better is it than GLM-5.2, really? +
Z.ai claims +50% over 5.2 on its own Z.ai Code Bench, and reports Terminal-Bench 3.0 rising from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5. Every one of those is a vendor-run figure on a vendor page — we carry them because they are what exists, not because they are neutral. No third-party evaluation of GLM-5.3 was available at the time of writing.
Q · 05 What is the cached-input storage charge? +
Z.ai marks it "Limited-time Free" across the whole GLM line, GLM-5.3 included. It is a separate charge from the per-token rates — you pay for keeping a cache alive, not for reading it — and we do not model it, because it is free today and Z.ai publishes no rate for when it stops being free.
Q · 06 What are the model's limits? +
1M context and 128K maximum output, text in and text out. Thinking mode, streaming, function calling, context caching, structured output and MCP are all supported per Z.ai's model page.