Last verified
FRONTIER MoE256K CONTEXTTEXTOPEN WEIGHTPROMPT CACHING

Hunyuan Hy3 API Pricing

Hunyuan Hy3 is Tencent's generally available Hunyuan flagship - a hybrid fast/slow-thinking 295B/21B MoE released July 6, 2026. Tencent Cloud International lists it at $0.132/M input, $0.528/M output, and $0.033/M cached input. Those are Tencent's own USD figures; the CN-region card prices the same model in yuan at Y1.00 / Y4.00 / Y0.25, and the two lists are published independently rather than converted.

Input - per 1M tokens
$0.132/M
CN card Y1.00/M CN list
Output - per 1M tokens
$0.528/M
CN card Y4.00/M CN list
Cached input - 75% off
$0.033/M
CN card Y0.25/M -75%
Effective - agentic blend
$0.089/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with Hunyuan Hy3's current Tencent Cloud International rate. Tweak the workload mix and cache hit rate; share the URL to share the calculation.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (tencent/Hunyuan-A13B-Instruct, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 873
Words 138
Tokens (estimated) 167 tokens
Cost as input · uncached $0.00002 USD
Cost as output · uncached $0.00009 USD
Cost as cached input $0.00001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Hunyuan Hy3 Current $0.132 cache $0.033 $0.528 $0.089 agentic 92/8 256K Tencent's GA flagship - agentic coding
Hunyuan Hy4 Preview $0.834 cache $0.042 $2.50 $0.37 the tier above, preview status 1M 770B model and 1M context when Hy3 runs out of room
Hunyuan 2.0 Think $0.591 $2.37 $0.733 pricier reasoning tier 128K Deep-thinking reasoning
Hunyuan 2.0 Instruct $0.473 $1.18 $0.53 pricier instruct tier 128K Broader Tencent instruction work
Hunyuan T1 $0.149 $0.595 $0.185 no cache discount documented elsewhere Legacy - off Tencent's price list since the TokenHub migration
Hunyuan TurboS $0.119 $0.298 $0.133 fast general tier documented elsewhere Legacy - off Tencent's price list since the TokenHub migration
Hunyuan A13B $0.0744 $0.298 $0.0923 similar blended cost 224K Low-cost Tencent traffic
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 cross-vendor budget benchmark 1M Cheap reasoning and coding

Frequently asked.

Practical Hunyuan Hy3 pricing questions, with Tencent's two separate price lists - International USD and CN-region CNY - kept apart from calculator assumptions.

Q · 01 What is Hunyuan Hy3 priced at today? +
Tencent Cloud International lists Hunyuan Hy3 at $0.132/M input, $0.528/M output, and $0.033/M cached input for the Singapore region. Tencent's CN-region card prices the same model at Y1.00 / Y4.00 / Y0.25, but that is a parallel list rather than the same number in another currency - converting it at today's reference rate of Y6.7209/USD gives $0.1488 / $0.5952 / $0.0372, about 13% above what Tencent actually charges. This page carried those converted figures until August 31, 2026; see corrections.
Q · 02 Does Hy3 have a cache-hit discount? +
Yes - $0.033/M on a cache hit, 75% below fresh input. This is a flat single-rate row with no input-length tiers. The separate Hy3 preview row was input-length tiered and was never this general-availability model; Tencent scheduled that preview row offline for August 31, 2026 while GA Hy3 continued unchanged.
Q · 03 Is Hunyuan Hy3 open weight? +
Yes. Tencent released Hy3 under the commercially friendly Apache 2.0 license on Hugging Face and ModelScope from day one, alongside the hosted API on Tencent Cloud TokenHub. It is a 295B-total / 21B-active Mixture-of-Experts model.
Q · 04 What is the context window? +
Hy3 supports a 256K-token context window and is a hybrid fast/slow-thinking model (interleaved thinking) aimed at coding agents, document automation, and multi-step tool calling.
Q · 05 How does Hy3 compare with Hy4 Preview? +
Hy4 Preview is a bigger model in a higher price band, not a replacement: 770B total parameters and 1M context against Hy3's 295B and 256K, at 6.3x the input rate. Because Hy4's cache-hit rate is only 1.3x Hy3's, the gap narrows to 4.2x on cache-heavy agentic traffic. Hy3 stays on the price list as the GA row and is still the cheaper choice for steady production work.
Q · 06 What happened to Hunyuan T1, TurboS and Lite? +
They are off Tencent's price list. As of August 31, 2026 none of the three appears on the billing page or the model list of the original Hunyuan platform, and none was carried over to TokenHub - that platform now heads its pages with a migration notice saying it will add no new models and has stopped supporting new purchases, while existing customers keep working. We hold them as legacy rather than retired because Tencent's API reference still names them and we have not confirmed that calls fail. Their pages keep the last published rate for the record.
Q · 07 Does this page include tax or enterprise discounts? +
No. AI//COST stores Tencent's public pre-tax list price and converts it to USD for cross-vendor comparison. TokenHub subscription plans, contract discounts, and region-specific billing are outside this page.