Last verified
2.4T TOTAL / 95B ACTIVE1M CONTEXTTHINKING INCLUDEDINTERNATIONAL LIST1M FREE TOKENS

Qwen3.8-2.4T-A95B API Pricing

Qwen3.8-2.4T-A95B is the parameter-named member of the Qwen3.8 family — 2.4T total parameters with 95B active. Alibaba's International (Singapore) rate card lists it at $2/M input and $6/M output in a single billing band up to 1M tokens, with Non-Thinking and Thinking modes priced identically. That is the same card as Qwen3.8 Max, so the two differ on architecture rather than cost. Alibaba also publishes a Global row at $1.65/$4.951 and a discounted batch row; the figures here are the International list, which is the tier this catalogue records.

Input - per 1M tokens
$2.00/M
Base token price standard
Output - per 1M tokens
$6.00/M
Output tokens standard
Cached input - not itemized
$2.00/M
Cache discount noted, price not itemized not itemized
Effective - agentic blend
$2.32/M
92/8 split - no cache discount
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with the International (Singapore) list rates. Thinking mode costs the same as non-thinking, so reasoning does not carry a separate rate — but it does raise the output token count. New accounts get 1 million free tokens, valid 90 days.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (Qwen/Qwen3.5-397B-A17B, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 440
Words 69
Tokens (estimated) 84 tokens
Cost as input · uncached $0.00017 USD
Cost as output · uncached $0.0005 USD
Cost as cached input $0.00017 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Qwen3.8-2.4T-A95B Current $2.00 $6.00 $2.32 2.4T / 95B active 1M Frontier Qwen with thinking
Qwen3.8 Max $2.00 $6.00 $2.32 same price, Max tier 1M The Max-tier sibling
Qwen3.7 Max $2.50 $7.50 $2.90 pricier 1M Previous Max flagship, text only
Qwen3.7 Plus $0.40 $1.60 $0.496 cheaper 1M Cheaper Qwen tier for bulk work
GPT-5.6 Terra $2.00 cache $0.20 $12.00 $1.52 cheaper 1.05M Same input rate, real cache discount
Claude Sonnet 4.6 $3.00 cache $0.30 $15.00 $2.05 cheaper 1M Long agentic coding sessions
DeepSeek V4 Pro $1.32 cache $0.044 $3.96 $0.569 cheaper 1M Cheapest 1M reasoning tier
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

Qwen3.8-2.4T-A95B pricing questions, with the International list kept apart from Alibaba's other rows.

Q · 01 How much does Qwen3.8-2.4T-A95B cost? +
$2/M input and $6/M output on Alibaba's International (Singapore) list, in a single billing band for prompts up to 1M tokens. Non-Thinking and Thinking modes are priced identically.
Q · 02 Why do other sites quote a lower price? +
Because Alibaba publishes several rows for the same model. Alongside the International list there is a Global row at $1.65 / $4.951 and a discounted batch row. This catalogue records the International list throughout, so figures stay comparable between Qwen models — mixing the tiers is the single easiest way to misread Alibaba's page.
Q · 03 How does it differ from Qwen3.8 Max? +
On price, not at all — both are $2 / $6 on the same card with the same 1M context. The difference is the architecture the name spells out: 2.4T total parameters with 95B active per token. Alibaba publishes no head-to-head benchmark between the two, and we do not carry one.
Q · 04 Is thinking mode billed differently? +
No. The rate card prices Non-Thinking and Thinking modes the same, so there is no separate reasoning rate to budget for. Thinking still costs more in practice, because chain-of-thought tokens are billed as output.
Q · 05 Is there a free allowance? +
Yes — 1 million free tokens, valid for 90 days from Model Studio activation, model release, or application approval, whichever is later. It is enough to size a workload, not to run one.