Last verified
FLAT TO 1MREADS IMAGESNO BATCH ROWINTERNATIONAL LIST1M FREE TOKENS

Qwen3.8-Flash API Pricing

Qwen3.8-Flash lists at $0.15/M input and $0.47/M output on Alibaba's International (Singapore) card. The number that matters is not the rate but the single billing band: one price from the first token to the millionth. Every other current Qwen-Flash charges by prompt length — Qwen3.7-Flash starts at $0.03/M but climbs to $0.20/M past 256K tokens. So the two models swap places at exactly that band edge, and neither is simply "cheaper" than the other.

Input - per 1M tokens
$0.15/M
Base token price flat to 1M
Output - per 1M tokens
$0.47/M
Output tokens flat to 1M
Cached input - not itemized
$0.15/M
Cache discount noted, price not itemized not itemized
Effective - agentic blend
$0.176/M
92/8 split - no cache discount
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with the International (Singapore) list rates. There is only one band to load: unlike its predecessors, this model charges the same rate whether the prompt is 2,000 tokens or 900,000. Thinking mode is supported and carries no separate rate — it only raises the output token count. New accounts get 1 million free tokens, valid 90 days.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (Qwen/Qwen3.5-397B-A17B, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 451
Words 75
Tokens (estimated) 86 tokens
Cost as input · uncached $0.00001 USD
Cost as output · uncached $0.00004 USD
Cost as cached input $0.00001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Qwen3.8-Flash Current $0.15 $0.47 $0.176 flat band 1M One rate across the whole window
Qwen3.7-Flash $0.03 $0.13 $0.038 short prompts only 1M Cheaper below 256K, dearer above
Qwen3.5-Flash $0.10 $0.40 $0.124 cheaper 1M The last flat-band Flash before this one
Qwen3.6-Flash $0.25 cache $0.25 $1.50 $0.35 pricier 1M Two bands, quadruples past 256K
Qwen3.8-27B $0.50 $3.00 $0.70 pricier 1M Open-weight sibling, same generation
GLM-5.3-Flash $0.15 cache $0.03 $0.50 $0.0875 cheaper 1M Same list input, half the blend
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 pricier 1M Five times the list rate, cache narrows it
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

Qwen3.8-Flash pricing questions, with the flat band kept apart from the tiered cards its predecessors use.

Q · 01 How much does Qwen3.8-Flash cost? +
$0.15/M input and $0.47/M output on Alibaba's International (Singapore) list, in a single billing band for prompts from 0 to 1M tokens. Thinking mode carries no separate rate.
Q · 02 Is it cheaper than Qwen3.7-Flash? +
It depends entirely on prompt length, and the crossover is sharp. Alibaba prices Qwen3.7-Flash in three steps by input tokens per request — $0.03/$0.13 up to 32K, $0.10/$0.40 to 256K, $0.20/$0.80 above that — while Qwen3.8-Flash charges $0.15/$0.47 throughout. Below 256K tokens the older model wins, and by a wide margin on short prompts. Above 256K the newer one is cheaper on both input and output. A 250K-token job costs about $0.029 on 3.7-Flash and $0.042 on 3.8-Flash; add ten thousand tokens to cross the band and the same job costs $0.060 against $0.044.
Q · 03 Does it accept images? +
Yes. Alibaba's visual-understanding documentation names Qwen3.8-Flash alongside Qwen3.8-Max and the Qwen3.7-Plus series as models that take up to 2,048 images per request — four times the 256-image ceiling the previous generation carries. Video input is documented for the Qwen3.8 series as a whole but not itemized for this model id, so it is not claimed here.
Q · 04 Is there a batch discount? +
Not on this model, and that is a break from the pattern. The rate card tags Qwen3.7-Flash, Qwen3.6-Flash, Qwen3.5-Flash and Qwen-Flash with a 50% batch inference discount badge; the Qwen3.8-Flash row carries only the context-caching note. No batch rate is published for it, so none is stored.
Q · 05 Why do other sites quote a lower price? +
Because Alibaba publishes several regional rows for the same model. Alongside the International list there is a China (Beijing) card at $0.113 / $0.382. This catalogue records the International list throughout, so figures stay comparable between Qwen models — mixing the regional cards is the single easiest way to misread Alibaba's page.
Q · 06 Is there a cache discount? +
The rate card marks context caching as discounted but publishes no per-model cached rate, so no cached figure is stored and the effective blend on this page assumes none. The cost of that silence is visible one row down the shelf: GLM-5.3-Flash lists at the same $0.15/M input, publishes a $0.03/M cache-hit rate, and lands at half the effective blend.
Q · 07 Is there a free allowance? +
Yes — 1 million free tokens, valid for 90 days from Model Studio activation, model release, or application approval, whichever is later. It is enough to size a workload, not to run one.