Qwen3.7 Flash API Pricing
Qwen3.7 Flash is Alibaba's cheapest current Qwen tier: $0.03/M input and $0.13/M output on the International rate card — about eight times under Qwen3.6 Flash. The catch is the tier ladder: that headline rate covers requests up to 32K tokens, and longer requests bill every token at the higher tier. Thinking and non-thinking cost the same, which the Max and Plus tiers do not offer. Pulled directly from alibabacloud.com daily.
Run the numbers.
Live calculator pre-loaded with the 0-32K Qwen3.7 Flash rates. If your requests run longer than 32K tokens, swap in the tier rate from the table below — Alibaba bills the whole request at the tier its input length falls into, not just the tokens above the threshold.
Real-world presets.
Consumer assistant turn
Ticket triage
Research planning turn
Document tagging at volume
Paste text. See tokens.
Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (Qwen/Qwen3.5-397B-A17B, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| Qwen3.7 Flash Current | $0.03 | $0.13 | $0.04 agentic 92/8 | 1M | Cheapest Qwen for short requests |
| Qwen3.6 Flash | $0.25 cache $0.25 | $1.50 | $0.35 the tier it replaces | 1M | Flat 0-256K entry tier |
| Qwen 3.5 Flash | $0.10 | $0.40 | $0.12 flat to 1M | 1M | No tier ladder to model |
| Qwen3 Coder Flash | $0.30 | $1.50 | $0.40 code specialist | 1M | Agentic coding on Qwen |
| Qwen3.7 Plus | $0.40 | $1.60 | $0.50 same generation, vision | 1M | Image and video understanding |
| Qwen3.7 Max | $2.50 | $7.50 | $2.90 Qwen frontier | 1M | Hardest Qwen reasoning |
| DeepSeek V4 Flash | $0.14 cache $0.0028 | $0.28 | $0.05 closest rival | 1M | Budget reasoning with real cache pricing |
| Gemini 3.5 Flash-Lite | $0.30 cache $0.03 | $2.50 | $0.27 Google budget tier | 1M | High-volume multimodal work |
Frequently asked.
Practical Qwen3.7 Flash pricing questions, with the tier ladder separated from the headline rate.
Q · 01 What is Qwen3.7 Flash priced at? +
$0.03/M input and $0.13/M output for requests of 32,000 tokens or fewer. Under AI//COST's 92/8 agentic blend the effective rate is $0.038/M, which makes it the cheapest Qwen tier currently published and one of the cheapest paid models in our catalogue.Q · 02 How does the tier ladder work? +
0-32K at $0.03 / $0.13, 32K-256K at $0.10 / $0.40, and 256K-1M at $0.20 / $0.80. One 40,000-token prompt therefore costs more than three 13,000-token prompts carrying the same content.Q · 03 What does a long-context request actually cost? +
32K-256K tier, so it bills at $0.10 / $0.40 and costs $0.016 per call. Priced naively at the headline 0-32K rate you would predict $0.0049 — a 3.3x understatement. The scenario cards above deliberately stay inside the 0-32K tier so the number on the card is the number you pay.Q · 04 Is it really cheaper than Qwen3.6 Flash? +
8x cheaper on input ($0.03 against $0.25), though 3.6 Flash's entry tier runs all the way to 256K while 3.7 Flash's stops at 32K. Compare like for like in the 32K-256K band and the gap is 2.5x ($0.10 against $0.25); at 256K-1M it is 5x ($0.20 against $1.00).Q · 05 Does thinking mode cost extra? +
Q · 06 What about cached input? +
Q · 07 Is there a batch discount? +
$0.015/M input and $0.065/M output on the entry tier. The discount applies to the tier rate, so a long batch request still climbs the ladder first.