Last verified
27B DENSEVISION ENCODER1M BILLING BANDINTERNATIONAL LIST1M FREE TOKENS

Qwen3.8-27B API Pricing

Qwen3.8-27B is the small, dense member of the Qwen3.8 family — and the one that reads images. Alibaba's International (Singapore) rate card lists it at $0.5/M input and $3/M output, a quarter of what the family's 2.4T-A95B and Max siblings cost. The size ordering and the capability ordering point opposite ways here: Qwen's own model card calls this one a Causal Language Model with Vision Encoder, while the far larger 2.4T-A95B is text-only.

Input - per 1M tokens
$0.50/M
Base token price standard
Output - per 1M tokens
$3.00/M
Output tokens standard
Cached input - not itemized
$0.50/M
Cache discount noted, price not itemized not itemized
Effective - agentic blend
$0.70/M
92/8 split - no cache discount
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with the International (Singapore) list rates. Non-Thinking and Thinking modes carry the same rate, so reasoning has no separate price — it only raises the output token count. New accounts get 1 million free tokens, valid 90 days.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (Qwen/Qwen3.5-397B-A17B, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 503
Words 83
Tokens (estimated) 96 tokens
Cost as input · uncached $0.00005 USD
Cost as output · uncached $0.00029 USD
Cost as cached input $0.00005 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Qwen3.8-27B Current $0.50 $3.00 $0.70 27B dense 1M Small Qwen3.8 that reads images
Qwen3.8-2.4T-A95B $2.00 $6.00 $2.32 pricier 1M The frontier sibling, text only
Qwen3.8 Max $2.00 $6.00 $2.32 pricier 1M Max tier with vision and video
Qwen3.6 Plus $0.50 $3.00 $0.70 same 1M Identical card, previous generation
Qwen3-VL Plus $0.20 $1.60 $0.312 cheaper 256K Cheaper vision Qwen, shorter context
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 cheaper 1M Dearer per token, cheaper with cache
Muse Spark 1.2 $1.25 cache $0.15 $4.25 $0.66 cheaper 1M 2.5x the list rate, cache closes it
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

Qwen3.8-27B pricing questions, with the International list kept apart from Alibaba's other rows.

Q · 01 How much does Qwen3.8-27B cost? +
$0.5/M input and $3/M output on Alibaba's International (Singapore) list, in a single billing band for prompts up to 1M tokens. Non-Thinking and Thinking modes are priced identically.
Q · 02 Does it accept images? +
Yes, and that is the surprise of this family. Qwen's model card describes Qwen3.8-27B as a Causal Language Model with Vision Encoder, and Hugging Face files it under image-text-to-text. The much larger Qwen3.8-2.4T-A95B is a plain causal language model — text only. Picking the bigger model to get vision would be the wrong move here.
Q · 03 Why do other sites quote a lower price? +
Because Alibaba publishes several rows for the same model. Alongside the International list there is a China (Beijing) card at $0.424 / $1.696. This catalogue records the International list throughout, so figures stay comparable between Qwen models — mixing the regional cards is the single easiest way to misread Alibaba's page.
Q · 04 What is the real context window? +
Qwen documents 262,144 tokens natively, extensible to 1,000,000. The 1M shown here is Alibaba's billing band — the rate card prices a single band from 0 to 1M input tokens per request, and the same basis is used for Qwen3.8 Max and 2.4T-A95B so the family stays comparable.
Q · 05 Is there a cache discount? +
The rate card marks context caching as discounted but publishes no per-model cached rate, so no cached figure is stored and the effective blend on this page assumes none. That matters against rivals: Gemini 3.7 Flash lists at 50% more per input token yet lands cheaper on an agentic blend, because its cached input rate is published at $0.075/M.
Q · 06 Is there a free allowance? +
Yes — 1 million free tokens, valid for 90 days from Model Studio activation, model release, or application approval, whichever is later. It is enough to size a workload, not to run one.