Last verified
ONLY PAY-PER-TOKEN GRANITE131K CONTEXTRESOURCE UNIT BILLINGNO CACHE RATEENTERPRISE PLATFORM

Granite 4 H Small API Pricing

Granite 4 H Small is the only Granite model IBM still sells per token. watsonx.ai lists $0.0000636 per 1,000 input tokens and $0.000265 per 1,000 output tokens, which is $0.0636/M and $0.265/M. Every other live Granite 4 model - the 3B, 8B and 30B rows, the vision model, tiny and micro - is sold by the GPU-hour instead, from $4.43/hour upward, which is not a token rate at all.

Input - per 1M tokens
$0.0636/M
IBM $0.0000636/1K per 1K x1000
Output - per 1M tokens
$0.265/M
IBM $0.000265/1K per 1K x1000
Cached input
n/a
No cache rate published N/A
Effective - agentic blend
$0.0797/M
92/8 split - no cache discount
§ 01 / TERMINAL

Run the numbers.

IBM meters inference in Resource Units, where one RU is 1,000 tokens and each model is assigned a pricing class rather than a bespoke rate. The figures here are those class rates scaled to a million tokens, so the arithmetic below matches an IBM invoice line.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · granite-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 600
Words 101
Tokens (estimated) 156 tokens
Cost as input · uncached $0.00001 USD
Cost as output · uncached $0.00004 USD
Cost as cached input $0.00001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Granite 4 H Small Current $0.0636 $0.265 $0.0797 no cache rate published 131K Cheap enterprise text work inside watsonx
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.682 cross-vendor cheap tier 200K High-volume support and classification
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 cross-vendor speed tier 1M Google's fast mainstream model
Mercury 2 $0.25 cache $0.025 $0.75 $0.12 cross-vendor low-latency row 128K Diffusion model sold on throughput
Qwen3.8-Flash $0.15 $0.47 $0.176 cross-vendor budget row 1M Flat-band Qwen light tier

Frequently asked.

What a Resource Unit is, why only one Granite model has a token price, and what happened to Granite 4.2.

Q · 01 What does Granite 4 H Small cost? +
IBM lists $0.0000636 per 1,000 input tokens and $0.000265 per 1,000 output tokens on the watsonx.ai supported-models table. Multiplied by 1,000 that is $0.0636/M input and $0.265/M output. Under this site's 92/8 blend the effective rate is $0.0797/M. IBM publishes no cache-hit rate for this model, so no cache discount is assumed here.
Q · 02 What is a Resource Unit? +
IBM's billing unit for foundation-model inference: 1 RU = 1,000 tokens. Rather than pricing each model individually, IBM assigns it a pricing class and charges the class rate per RU - this model is class 18 on input and class 5 on output. Divide your monthly tokens by 1,000, round up, multiply by the class rate. The per-million figures on this page are that same arithmetic, rescaled.
Q · 03 Do I need an expensive plan to get this rate? +
No. IBM's Essentials pay-as-you-go plan starts at USD 0/month, and the token rates apply there. The Standard plan starts at USD 1,110/month and buys enterprise production features, not a better token price. Reports of a four-figure monthly minimum to use Granite at all do not match IBM's own pricing page.
Q · 04 Why is this the only Granite model with a token price? +
Because IBM splits its catalogue in two. The supported-models table has a column for "Provided with watsonx.ai (Pay per token)" and another for "Deploy on demand (Pay by the hour)". Only granite-4-h-small sits in the first among live models - the other live Granite 4 rows, including the 3B, 8B and 30B variants, the vision model and the tiny and micro sizes, are in the hourly column, billed from $4.43/hour for a single L40S GPU. Two Granite rows do carry token prices but IBM marks them Deprecated.
Q · 05 What about Granite 4.2? +
It is not on IBM's pricing or supported-models pages at all. Granite is open-weight, so resellers can host a release before - or without - IBM selling it per token first-party; that is how it reached the model listings that flagged it to us. Until IBM publishes a rate, there is nothing here to quote, and a reseller's price is not the vendor's price.
Q · 06 Is the context window the same as everyone else measures? +
Read it carefully: IBM states 131,072 tokens as input plus output together. Most vendors quote an input window and a separate max-output figure. On like-for-like long-context work that makes IBM's number slightly smaller than it looks next to a nominally equal 128K context elsewhere.