Last verified
QWEN 3.8 FLAGSHIPOUT OF PREVIEW1M CONTEXTTEXT + VISION + VIDEOTHINKING SAME PRICE

Qwen3.8 Max API Pricing

Qwen3.8 Max is Alibaba's newest Qwen Max flagship, and as of August 2026 it finally has a price. The International / Singapore row lists $2/M input and $6/M output for 0-1M tokens, published directly in US dollars. That undercuts Qwen3.7 Max, the flagship it replaces, on both sides of the meter — the newer model is the cheaper one. Thinking and non-thinking modes bill at the same rate.

Input - per 1M tokens
$2.00/M
Base token price standard
Output - per 1M tokens
$6.00/M
Output tokens standard
Cached input - not itemized
$2.00/M
Cache discount noted, price not itemized not itemized
Effective - agentic blend
$2.32/M
92/8 split - no cache discount
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with the International Qwen3.8 Max rates. Alibaba publishes no per-model cache rate, so the cache slider stays off — what you see is the undiscounted cost.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (Qwen/Qwen3.5-397B-A17B, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 407
Words 64
Tokens (estimated) 78 tokens
Cost as input · uncached $0.00016 USD
Cost as output · uncached $0.00047 USD
Cost as cached input $0.00016 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Qwen3.8 Max Current $2.00 $6.00 $2.32 agentic 92/8 1M Newest Qwen flagship, now multimodal
Qwen3.7 Max $2.50 $7.50 $2.90 pricier 1M Previous Max flagship, text only
Qwen3.7 Plus $0.40 $1.60 $0.496 cheaper 1M Cheaper Qwen tier for bulk work
GPT-5.6 Terra $2.00 cache $0.20 $12.00 $1.52 cheaper 1.05M Same input rate, real cache discount
Claude Sonnet 4.6 $3.00 cache $0.30 $15.00 $2.05 cheaper 1M Long agentic coding sessions
DeepSeek V4 Pro $1.32 cache $0.044 $3.96 $0.569 cheaper 1M Cheapest 1M reasoning tier
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

Short answers for teams checking Qwen3.8 Max pricing, regional rate cards, and what changed when it left preview.

Q · 01 What is Qwen3.8 Max priced at? +
Alibaba's International (Singapore) rate card lists $2/M input and $6/M output, published directly in US dollars, in a single billing band covering 0-1M input tokens per request. Under AI//COST's 92/8 agentic blend the effective planning figure is $2.32/M — there is no cache discount to apply, so that is simply the undiscounted mix.
Q · 02 Why do I see $1.65 and $4.951 elsewhere? +
Because Alibaba publishes several rate cards on one page and they are different products. $1.65/$4.951 is the Global / China deployment card; $2/$6 is the International (Singapore) card. AI//COST quotes the International list for every Alibaba model, so the numbers stay comparable across our catalogue. Mixing the two is exactly what produced the pricing error we corrected on July 27, 2026.
Q · 03 Did the price change when it left preview? +
There was no price to change. From its WAIC announcement on July 19, 2026 the model was qwen3.8-max-preview, reachable only through a Token Plan subscription in the Beijing region, with no per-token rate published anywhere. We ran the page with a blank price board rather than estimate one. The formal qwen3.8-max ID has now replaced it and carries the first real rate this model has ever had.
Q · 04 Is it cheaper than Qwen3.7 Max? +
Yes, on both sides: $2 against $2.50 input, $6 against $7.50 output — about 25% cheaper on the blended agentic figure, and it adds image and video input that Qwen3.7 Max does not have. One caveat worth checking before you migrate: Qwen3.7 Max currently carries a limited-time 50% discount on its list price, so what you are billed today may be lower than its list suggests.
Q · 05 Is there a Batch discount? +
Not on the International card. Alibaba states that batch calls bill at 50% of the real-time price for models marked as supporting them, and the International row for Qwen3.8 Max carries no batch marker — only a context-caching one. The Global/China rows for this model do show the batch marker. We record no batch rate rather than assume the discount carries across regions.
Q · 06 Is context caching priced? +
It is offered but not itemized. The rate card marks Qwen3.8 Max as having a context-caching discount without publishing a per-model cache-hit rate, so we leave the cached row at the standard input price and keep the calculator's cache slider off. Every figure on this page is therefore the undiscounted cost — your real bill with heavy caching will be lower by an amount Alibaba does not disclose.
Q · 07 Which API model ID should I use? +
qwen3.8-max. The preview ID qwen3.8-max-preview has been withdrawn from Alibaba's rate cards. The model is served from Beijing, Singapore, Tokyo, Frankfurt and Virginia endpoints, and Alibaba exposes OpenAI-compatible, Anthropic-compatible and DashScope interfaces for it.
Q · 08 Does it accept images and video? +
Yes. Alibaba lists Qwen3.8 Max under both text generation and image/video understanding in its model catalogue, which is a change from Qwen3.7 Max — that tier is text-only. The rate card publishes a single token price covering the model rather than separate per-modality rates.