AI//COST
Pricing
All models & pricing →
Alibaba (Qwen) 33 Zhipu (Z.ai / GLM) 16 ByteDance (Doubao) 14 OpenAI 13 Anthropic 10 Baichuan 10 Tencent (Hunyuan) 9 Google 8 MiniMax 8 Mistral AI 8 SpaceXAI 7 Cohere 4 Moonshot (Kimi) 4 Perplexity 4 NVIDIA 3 Reka 3 AI21 Labs 2 Baidu 2 DeepSeek 2 Sakana AI 2 Upstage 2 Amazon 1
CalculatorCompareAgentsJournalMethodologyAboutContact
Pricing
Alibaba (Qwen) 33Zhipu (Z.ai / GLM) 16ByteDance (Doubao) 14OpenAI 13Anthropic 10Baichuan 10Tencent (Hunyuan) 9Google 8MiniMax 8Mistral AI 8SpaceXAI 7Cohere 4Moonshot (Kimi) 4Perplexity 4NVIDIA 3Reka 3AI21 Labs 2Baidu 2DeepSeek 2Sakana AI 2Upstage 2Amazon 1
CalculatorCompareAgentsJournalMethodologyAboutContact
↑↓ navigate ↵ open esc close See all results →
Home / Corrections
Corrections Log PUBLIC UPDATED ON CHANGE

Corrections log

When we get something wrong, we say so publicly. Every correction is logged with date, page, what was wrong, what's correct, and who flagged it. See the corrections process in the methodology.

§01

Open log

# log

The corrections log records material and severe corrections per the methodology severity policy. Minor fixes (typos, broken links) are not logged separately.

Aug 31, 2026 material /llm-cost/tencent/hunyuan-hy3/ and the Tencent hub - every published rate for Hunyuan Hy3

WasHy3 was listed at $0.141/M input, $0.563/M output and $0.035/M cached input. Those figures were not Tencent's: we produced them by converting the CN-region CNY card (Y1.00 / Y4.00 / Y0.25) at the Y7.10 rate this catalogue used until that day. Tencent Cloud International publishes its own USD price list for the same model at $0.132 / $0.528 / $0.033, so every rate we showed was about 6.8% too high, and so was the effective blended rate and each of the four scenario costs built on it. A reader budgeting from this page would have overstated their Hy3 spend.

NowAll Hy3 rates now come from Tencent Cloud International's USD list, the same source as the new Hy4 Preview row. The two lists are not one number in two currencies: Hy3's USD implies about Y7.58 to the dollar and Hy4 Preview's implies Y7.19, so no single conversion reproduces them and the published USD has to be read rather than derived. Tencent's prices did not change - the price history on that page is deliberately flat, because what moved was our reading of it. This is the same failure we recorded for Alibaba in July, when converting a CN rate card put fourteen prices above the vendor's own International list; the rule that came out of it had not been carried across to Tencent.

Flagged by Internal audit.

Aug 27, 2026 material /blog/claude-code-cache-misses-usage-limits/ - the claim that a subscription has no control over the subagent cache TTL

WasThe post, published 2026-08-23, said that subagents get the five-minute cache TTL even on a subscription "with no flag on a subscription to widen it", and presented ENABLE_PROMPT_CACHING_1H and FORCE_PROMPT_CACHING_5M as the only TTL controls. Anthropic has since split Claude Code's requests into two TTL buckets - the main conversation, and everything else (subagents, workflows, in-process teammates, forks, compaction, session titles) - and given each bucket its own control. A reader running fan-outs on a subscription would have concluded that the shortest cache lifetime in the product was not adjustable, when it is.

NowThe post now describes both buckets and all four controls: promptCacheTtl and CLAUDE_CODE_PROMPT_CACHE_TTL for the main conversation, subagentPromptCacheTtl and CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL for everything else, each taking 5m or 1h, with FORCE_PROMPT_CACHING_5M overriding both and ENABLE_PROMPT_CACHING_1H now below both per-bucket controls in the precedence order. Anthropic's documentation puts the requirement at Claude Code v2.1.242 or later; the vendor's changelog lists both settings as landing in v2.1.243, after the post was published. No price, calculation or table was affected - the arithmetic in that post is unchanged and still checks against the snapshot. Caught while researching the follow-up piece on parallel agents, which re-read the same documentation page four days later and found it restructured.

Flagged by Internal audit.

Aug 27, 2026 material /llm-cost/alibaba/ - the input modalities of Qwen3.7-Flash and Qwen3.6-Plus

WasBoth models were recorded as text-only, so the provider hub's lineup printed "text-only" beside each of them. Alibaba documents both as accepting image input: its Model Studio visual-understanding page names Qwen3.7-Flash and Qwen3.6-Plus among the models that take up to 256 images per request. The field is not cosmetic - it is the only signal on the hub that tells a reader whether a model can be sent a screenshot, and a team filtering for vision would have skipped two models that do it.

NowBoth now record text, vision and code. Per-token prices were correct throughout and are unchanged; no calculator, scenario or effective-rate figure was affected. Found by the control step of our model-intake process, which re-reads a vendor's whole card when a new model from that vendor is added - the same step that caught the Grok modality error in August. This time the check was widened from prices to the vendor's capability documentation, because a price has a shelf of gates behind it and a modality has none.

Flagged by Internal audit.

Aug 23, 2026 material /llm-cost/deepseek/ and both V4 model pages - the prices and answers left behind by the move to time-of-day billing

WasDeepSeek replaced its flat rate card with peak/off-peak billing on August 16, 2026. The stored rates were updated that day and every calculator, tile and scenario followed, but six pieces of prose did not. The provider hub's headline statistic read "Cheapest tier $0.14/M" for a model that now costs $0.44 at peak and $0.22 off-peak; a hub answer compared V4 Pro to GPT-5.5 using $0.87 output when the rate was $3.96, and quoted cache-hit input at $0.0036 instead of $0.044; the hub answered the question "Is there an off-peak discount?" with "No discount" on the very day off-peak became the billing model; and the V4 Flash page still said we had not pre-applied the change. Two further answers were stale in a different direction: both model pages said the Responses API was unsupported on V4 Pro with support "coming in early August 2026", and both repeated a vendor warning about an unspecified future price rise that had already arrived and been removed from the page.

NowAll six now state the live card: $0.44/$0.014/$1.32 per million at peak on V4 Flash, $1.32/$0.044/$3.96 on V4 Pro, and exactly half of each off-peak, which covers 17 hours of a weekday and all weekend. The Responses API is supported on all three DeepSeek models as of today; the remaining difference is concurrency, 500 on Pro against 2500 on Flash. No stored price was ever wrong and no calculator was affected - the defect was confined to hand-written prose and to one hand-written statistic. To stop that statistic going stale again, the hub's model count, cheapest tier and context window are now computed from the snapshot at build time rather than typed in.

Flagged by Internal audit.

Aug 13, 2026 material /llm-cost/xai/grok-4-20-multi-agent-0309/, /llm-cost/xai/, and the sibling tables on Grok 4.3 and Grok 4.5 — context window and input modalities

WasGrok 4.20 Multi-Agent was published with a 2M-token context window and described as "the only Grok variant with 2M context" — a claim repeated in the model page subtitle and hero tag, the provider hub's headline "Max context" statistic, a use-case pick for "the longest context window in the Grok line", a competitive takeaway against Gemini 2.5 Pro, and the sibling tables on two other Grok pages. SpaceXAI has never published that figure. The same rows also recorded the model as text-only, along with Grok 4.3 and both dated Grok 4.20 snapshots.

NowThe window is 1M. All three SpaceXAI surfaces agree — the pricing table, the model page spec block and the raw page source all state 1,000,000 tokens, and the page contains no 2,000,000 anywhere. Input modalities are text + image for the whole Grok 4.20 family and for Grok 4.3, not text-only. Per-token prices were correct throughout and are unchanged, so calculators, cost scenarios and effective-rate figures were never affected; the error was confined to the context and modality claims. Found by the routine control step in our model-intake process, which re-reads every existing row for a vendor whenever a new model from that vendor is added.

Flagged by Internal audit.

Aug 13, 2026 material /llm-cost/anthropic/claude-sonnet-5/, /llm-cost/anthropic/, /compare/claude-sonnet-vs-haiku/ and two blog posts — the cancelled September price rise

WasEight surfaces stated that Claude Sonnet 5 would move from $2/$10 to $3/$15 per million tokens on September 1, 2026, when its introductory pricing expired. That was true when we published it — Anthropic's pricing page carried the schedule and the September rates. We built on it: the Sonnet-vs-Haiku comparison awarded Haiku a 'price stability' verdict on the strength of the coming rise, one pricing row was devoted to the September rate card, and two blog posts used the date as a worked example of a bill changing without a code change.

NowAnthropic cancelled the increase. Its pricing page now states that $2/$10 'is now the standard price' and that 'the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.' The stored rate was always $2/$10 and never wrong; what was wrong was the forecast attached to it. The scheduled-change block, the pre-computed September rates and the verdict built on them have been removed, and both blog posts carry a dated update rather than a silent edit.

Flagged by Internal audit.

Aug 2, 2026 material /llm-cost/openai/ — GPT-5.6 Sol, Terra and Luna long-context rates

WasAll three GPT-5.6 pages published long-context rates at half the real price. OpenAI's pricing page repeats the same table once per service tier, and at intake on July 10 we read the Batch table's long-context columns as if they were Standard. Sol showed $5/$0.50/$22.50 against a real $10/$1/$45, and Terra and Luna were understated the same way. The figures appeared both in a quote-board footnote and in a dedicated FAQ answer on each page, so a reader sizing a >272K-token request would have budgeted half of what OpenAI charges.

NowLong-context rates are now Sol $10/$1/$45, Terra $4/$0.40/$18 and Luna $0.40/$0.04/$1.80 — the vendor's long-context tier is double the standard input and cached-input rate and 1.5x the standard output rate, which every other OpenAI model in our catalogue already followed. The short-context rates that drive the calculators and every comparison were correct throughout; only the long-context tier was wrong.

Flagged by Internal audit.

Aug 2, 2026 material Site-wide — displayed per-token rates

WasPrice displays rounded every rate to two decimals, which is too coarse for a catalogue spanning $0.01 to $150 per million tokens. 115 of 770 displayed rates sat more than 1% from the real figure and 21 were more than 5% out — a $0.015 rate read as "$0.01", 33% low. Separately, provider-hub lineup cards rounded anything above $1 to whole dollars, so 198 rates across 132 models were wrong on those cards: Gemini 3.6 Flash's real $1.50 input showed as "$2". Errors ran in both directions, and the visible price disagreed with the exact figure we publish in structured data.

NowRates now render to three significant figures with a two-decimal floor, so $0.435 reads as $0.435 and $1.50 reads as $1.50 while $15.00 keeps its column alignment. Worst-case display error drops from 33% to 0.38% and nothing exceeds 1%. Whole-dollar rounding on lineup cards is gone. The pricing snapshot, calculators and structured data always held the exact values — only the rendering was lossy — and 38 authored page figures that had been written pre-rounded to two decimals were realigned to the snapshot at the same time.

Flagged by Internal audit.

Aug 2, 2026 material /llm-cost/deepseek/ — provider hub

WasThe DeepSeek hub reported a price change that never happened. A dated timeline entry, the V4 Pro lineup card, the lab background and an FAQ all stated that the 75% launch promo ended on May 31, 2026 and that V4 Pro had reverted to $1.74/M input and $3.48/M output — four times the real price. DeepSeek announced that expiry but never applied it. The reversion was written up in advance as a settled fact, so when the vendor let the date pass, the page kept asserting a list price the vendor had never charged.

NowV4 Pro lists at $0.435/M input and $0.87/M output, re-verified against DeepSeek's pricing page on August 2, 2026. The hub now records the expiry date passing without a reprice and treats the discounted figure as the standard rate. The underlying pricing snapshot always held the correct rates, so calculators and model pages were unaffected — the wrong numbers were confined to the hub's prose.

Flagged by Internal audit.

Jul 27, 2026 material /compare/windsurf-vs-cursor/

WasThe comparison stated that Cursor has no JetBrains support at all — "Cursor has none", a ✗ in the feature matrix, and a whole scenario advising JetBrains teams away from it on that basis.

NowCursor does reach JetBrains, through the Agent Client Protocol: the agent reads files and runs terminal commands inside IntelliJ, PyCharm and WebStorm. It is weaker than a native plugin — it needs a paid Cursor plan plus the JetBrains AI Assistant plugin 2025.1+, and Tab completion stays in the Cursor app — so Windsurf still wins that row, but the page now says what the limitation actually is instead of denying the feature exists.

Flagged by Internal audit.

Jul 27, 2026 material /llm-cost/alibaba/ — 14 Qwen model pages, and every page quoting their rates

WasFourteen Qwen models carried prices converted from Alibaba's China-region CNY rate card — Qwen3.7 Max, for example, at $2.7707/M input and $8.3119/M output. Alibaba publishes its own International (Singapore) list directly in US dollars, and those figures are lower, so every affected rate, calculator result, cost scenario, sibling shelf and comparison page overstated the real price by roughly 8-11%.

NowAll fourteen now carry the vendor's published International USD list, read from the English pricing page: Qwen3.7 Max is $2.50/M input and $7.50/M output. Calculators, price tiles, cost scenarios, price history and every sibling shelf were recomputed to match, and the currency-conversion fields were removed so nothing implies a derived figure. The model-intake note now states that the CN rate card must not be converted when an International USD list exists.

Flagged by Internal audit.

Jul 20, 2026 severe /methodology/ (§ agent testing, § benchmark handling)

WasThe methodology claimed we run all 500 SWE-bench Verified tasks ourselves on a clean box at ~$170–500 per agent, and that our test harness and result spreadsheets are public. We do none of that — every benchmark score on the site is read off a third-party leaderboard.

NowBoth sections were retracted and rewritten to state plainly that we do not run benchmarks ourselves, that each score carries the leaderboard it came from and the date we read it, and that the figures come from neutral third-party harnesses. The false per-agent cost table and 'reproduced by us' claims are gone.

Flagged by Internal audit.

Jul 20, 2026 material /llm-cost/ — model pricing pages

WasEffective and raw rates in the sibling-comparison tables had gone stale after vendor repricings that never propagated: Grok 4.3 overstated 142%, Gemini 3.5 Flash understated roughly 6×, Mistral Small 4 wrong on 30 pages, and Claude Mythos 5 shown at $13.20 against a real $6.82. The same model could read one effective rate on its own page and a different one in a comparison (Opus 4.8: $3.21 vs $3.41).

NowAll effective-blend and raw rates are now derived from the central pricing snapshot at build time using one shared formula, so a model reads the same number everywhere and cannot drift again. 927 stored values across 188 files were corrected, and a publish gate now fails the build on any disagreement.

Flagged by Internal audit.

Jul 20, 2026 material /compare/claude-vs-chatgpt/

WasThe cache-write premium (1.25× input) was recorded only for Anthropic models, so the effective-cost comparison charged Claude for a cost it did not charge GPT — which levies the same premium on GPT-5.6. The page overstated the Claude-vs-GPT cost gap by about 6%.

NowCache-write billing was audited across 16 providers and the GPT-5.6 rates filled in, so both sides are charged on the same basis. Where a vendor genuinely publishes no cache-write rate, that absence is now recorded explicitly rather than read as zero.

Flagged by Internal audit.

Jul 20, 2026 material Site-wide — price displays and Product schema

WasSub-cent rates such as DeepSeek V4 Flash cached input ($0.0028/M) rendered as '$0.00' across price tables — on a price desk that reads as free — and, worse, were published to Google in schema.org offers as a $0.00 (free) price.

NowA shared sub-cent formatter now shows the real figure to four significant digits, genuine free tiers read 'Free', and schema prices escalate precision so no real rate serialises as $0.00. Zero all-zero dollar strings remain in the build.

Flagged by Internal audit.

Jul 17, 2026 material /compare/ — comparison pages

WasGoogle Search Console flagged merchant-listing errors on the comparison pages: the structured-data offers were malformed, so the pages' pricing could be misrepresented in Google's shopping surfaces.

NowThe offer markup was corrected to a valid digital-product shape across the comparison pages and the fix verified live; the listings validate cleanly.

Flagged by Google Search Console.

★ Spot an error?

If you see something wrong — incorrect pricing, outdated benchmark score, factually wrong claim — email [email protected] with the page URL, the incorrect claim, and a source for the correct information. We review within 48 hours and either correct or push back with reasoning.

§02

Severity levels

# severity

Three tiers, defined in the methodology:

  • Minor — typo, broken link, formatting glitch. Fixed silently with a timestamp on the affected page. Not logged here.
  • Material — incorrect price, wrong benchmark score, wrong feature claim. Page corrected + entry in this log + reply to whoever flagged it.
  • Severe — incorrect ranking, retracted claim. Page corrected + entry here + apology paragraph in next quarterly editorial review.
§03

What we do not silently change

# policy

We do not:

  • Silently edit material claims after publication.
  • Gaslight readers about what we previously published.
  • Delete corrections from this log once issued.
  • Issue a correction without naming the page affected.

We do:

  • Issue corrections within 48 hours of a material error being identified.
  • Keep this log indexable + linkable per the methodology.
  • Cite the corrected source in our methodology version log when a correction triggers a methodology update.
  • Reply to the person who flagged the error with confirmation.
AI//COST

The independent price desk for every AI model on the market.

Affiliate disclosure · Some links are tracked. Editorial rankings are never paid. Sponsored slots are visually marked. See full disclosure.

Tools

  • Calculators
  • Comparisons
  • Pricing hub
  • Rankings · soon
  • Integrations · soon

Research

  • AI Agents
  • Benchmarks · soon
  • Capability Q&A
  • Case studies · soon
  • Glossary · soon

Editorial

  • The Journal
  • Guides · soon
  • The desk · soon
  • Methodology
  • Corrections

Company

  • About
  • Contact
  • Press · soon
  • Privacy
  • Terms
© 2026 AI//COST · Vol. 02, Issue 047 · Published continuously since 2025
Search Methodology Privacy Terms Contact