Granite 4 H Small API Pricing
Granite 4 H Small is the only Granite model IBM still sells per token. watsonx.ai lists $0.0000636 per 1,000 input tokens and $0.000265 per 1,000 output tokens, which is $0.0636/M and $0.265/M. Every other live Granite 4 model - the 3B, 8B and 30B rows, the vision model, tiny and micro - is sold by the GPU-hour instead, from $4.43/hour upward, which is not a token rate at all.
Run the numbers.
IBM meters inference in Resource Units, where one RU is 1,000 tokens and each model is assigned a pricing class rather than a bespoke rate. The figures here are those class rates scaled to a million tokens, so the arithmetic below matches an IBM invoice line.
Real-world presets.
Internal assistant turn
Knowledge-base lookup
Document triage
Intent routing
Paste text. See tokens. See cost.
This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| Granite 4 H Small Current | $0.0636 | $0.265 | $0.0797 no cache rate published | 131K | Cheap enterprise text work inside watsonx |
| Claude Haiku 4.5 | $1.00 cache $0.10 | $5.00 | $0.682 cross-vendor cheap tier | 200K | High-volume support and classification |
| Gemini 3.7 Flash | $0.75 cache $0.075 | $3.75 | $0.481 cross-vendor speed tier | 1M | Google's fast mainstream model |
| Mercury 2 | $0.25 cache $0.025 | $0.75 | $0.12 cross-vendor low-latency row | 128K | Diffusion model sold on throughput |
| Qwen3.8-Flash | $0.15 | $0.47 | $0.176 cross-vendor budget row | 1M | Flat-band Qwen light tier |
Frequently asked.
What a Resource Unit is, why only one Granite model has a token price, and what happened to Granite 4.2.
Q · 01 What does Granite 4 H Small cost? +
$0.0000636 per 1,000 input tokens and $0.000265 per 1,000 output tokens on the watsonx.ai supported-models table. Multiplied by 1,000 that is $0.0636/M input and $0.265/M output. Under this site's 92/8 blend the effective rate is $0.0797/M. IBM publishes no cache-hit rate for this model, so no cache discount is assumed here.Q · 02 What is a Resource Unit? +
Q · 03 Do I need an expensive plan to get this rate? +
Q · 04 Why is this the only Granite model with a token price? +
granite-4-h-small sits in the first among live models - the other live Granite 4 rows, including the 3B, 8B and 30B variants, the vision model and the tiny and micro sizes, are in the hourly column, billed from $4.43/hour for a single L40S GPU. Two Granite rows do carry token prices but IBM marks them Deprecated.