Last verified
CHEAPEST QWEN1M CONTEXTTIERED BY LENGTHTHINKING SAME PRICE50% BATCH

Qwen3.7 Flash API Pricing

Qwen3.7 Flash is Alibaba's cheapest current Qwen tier: $0.03/M input and $0.13/M output on the International rate card — about eight times under Qwen3.6 Flash. The catch is the tier ladder: that headline rate covers requests up to 32K tokens, and longer requests bill every token at the higher tier. Thinking and non-thinking cost the same, which the Max and Plus tiers do not offer. Pulled directly from alibabacloud.com daily.

Input - per 1M tokens
$0.03/M
International 0-32K tier -88%
Output - per 1M tokens
$0.13/M
Thinking billed the same -91%
Long-context tier - 256K to 1M
$0.20/M
Pairs with $0.80/M out 6.7x
Effective - agentic blend
$0.04/M
92/8 split - no cache discount
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with the 0-32K Qwen3.7 Flash rates. If your requests run longer than 32K tokens, swap in the tier rate from the table below — Alibaba bills the whole request at the tier its input length falls into, not just the tokens above the threshold.

$ /mo
Workload split
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (Qwen/Qwen3.5-397B-A17B, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 619
Words 100
Tokens (estimated) 118 tokens
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Qwen3.7 Flash Current $0.03 $0.13 $0.04 agentic 92/8 1M Cheapest Qwen for short requests
Qwen3.6 Flash $0.25 cache $0.25 $1.50 $0.35 the tier it replaces 1M Flat 0-256K entry tier
Qwen 3.5 Flash $0.10 $0.40 $0.12 flat to 1M 1M No tier ladder to model
Qwen3 Coder Flash $0.30 $1.50 $0.40 code specialist 1M Agentic coding on Qwen
Qwen3.7 Plus $0.40 $1.60 $0.50 same generation, vision 1M Image and video understanding
Qwen3.7 Max $2.50 $7.50 $2.90 Qwen frontier 1M Hardest Qwen reasoning
DeepSeek V4 Flash $0.14 cache $0.0028 $0.28 $0.05 closest rival 1M Budget reasoning with real cache pricing
Gemini 3.5 Flash-Lite $0.30 cache $0.03 $2.50 $0.27 Google budget tier 1M High-volume multimodal work

Frequently asked.

Practical Qwen3.7 Flash pricing questions, with the tier ladder separated from the headline rate.

Q · 01 What is Qwen3.7 Flash priced at? +
Alibaba's International (Singapore) rate card lists $0.03/M input and $0.13/M output for requests of 32,000 tokens or fewer. Under AI//COST's 92/8 agentic blend the effective rate is $0.038/M, which makes it the cheapest Qwen tier currently published and one of the cheapest paid models in our catalogue.
Q · 02 How does the tier ladder work? +
Alibaba prices by total input length per request, and the whole request bills at the tier it lands in — not just the tokens above the threshold. For Qwen3.7 Flash: 0-32K at $0.03 / $0.13, 32K-256K at $0.10 / $0.40, and 256K-1M at $0.20 / $0.80. One 40,000-token prompt therefore costs more than three 13,000-token prompts carrying the same content.
Q · 03 What does a long-context request actually cost? +
Take a 120,000-token input with 10,000 tokens out. That falls in the 32K-256K tier, so it bills at $0.10 / $0.40 and costs $0.016 per call. Priced naively at the headline 0-32K rate you would predict $0.0049 — a 3.3x understatement. The scenario cards above deliberately stay inside the 0-32K tier so the number on the card is the number you pay.
Q · 04 Is it really cheaper than Qwen3.6 Flash? +
Yes, at every tier, but by different margins. On the entry tier it is roughly 8x cheaper on input ($0.03 against $0.25), though 3.6 Flash's entry tier runs all the way to 256K while 3.7 Flash's stops at 32K. Compare like for like in the 32K-256K band and the gap is 2.5x ($0.10 against $0.25); at 256K-1M it is 5x ($0.20 against $1.00).
Q · 05 Does thinking mode cost extra? +
No. Alibaba's price table marks this model "Non-Thinking and Thinking modes" on a single row, so reasoning tokens bill at the same output rate. That is not true across the range: the Max and Plus tiers price thinking output separately, so a thinking workload does not scale linearly when you move up the Qwen ladder.
Q · 06 What about cached input? +
The model is marked as eligible for Alibaba's context cache discount, but the vendor publishes no per-model cache rate on the pricing page — the documented rule is that explicit-cache writes bill at 125% of standard input and hits at 10%. We record only published per-model figures, so this page carries no cache rate and the blended number above assumes no cache discount. Check the Context Cache docs before modelling savings.
Q · 07 Is there a batch discount? +
Yes. Alibaba marks Qwen3.7 Flash with a 50% batch inference discount, which puts batch at $0.015/M input and $0.065/M output on the entry tier. The discount applies to the tier rate, so a long batch request still climbs the ladder first.
Q · 08 Does it accept images? +
Not per Alibaba's own capability overview, which lists qwen3.7-flash under text generation only. Qwen3.7 Plus is the same-generation model that also appears under image and video understanding, so route multimodal work there rather than assuming the Flash tier inherits it.