Last verified
PUBLISHING PRINCIPLES METHODOLOGY v1.3 REVIEWED QUARTERLY PUBLIC + AUDITABLE

How we work

The operating manual for everything we publish. How pricing data gets into the snapshot, where benchmark scores come from, how rankings are scored, how head-to-heads are built. Written so anyone can audit our claims — and challenge them when we're wrong.

Methodology version
v1.3
Reviews per year
4
Cool-off pass
48h
Verification
Manual
Corrections log
§ 01

Editorial principles.

#principles

Six rules govern everything we publish. They are locked into the methodology — changing them requires a quarterly review with a written rationale, not an editorial decision. The rules below are the surface; the rest of this document is how they're operationalized.

01

Verify before asserting

Anything that is not a well-known constant must be verified at the time of publication. Prices change, APIs deprecate, benchmark scores update. We don't quote from memory or training data. Each claim has a timestamp and a source.

02

Score before ranking

Rubrics are locked before testing. We define criteria + weights publicly, then run tests. Scores for one candidate are not visible while another is being scored — each pass is blind to the others. Rankings are a function of scores, not preferences.

03

Honesty over diplomacy

If a tool is bad at something, we say so. Every brand profile has a "Don't use it for" section. If we'd recommend a competitor for a use case, we link to the competitor — even when our affiliate relationship is with the subject.

04

Sponsorship quarantined

Editorial rankings are not for sale. Sponsored slots exist on some pages but are visually quarantined (dashed warm border, "ⓘ SPONSORED" tag). They cannot influence editorial ranking, scoring, or "best of" picks.

05

Corrections are public

When we get something wrong, we log it openly. The corrections page lists every correction with date, page, what was wrong, and what's now correct. We don't silently edit. We don't gaslight readers about what we used to say.

06

Methodology is versioned

This document is v1.3. Major changes ship as version bumps, each logged with a date and a summary at the foot of this page — so you can see exactly what shifted in our approach and when. No silent rewrites of how we work.

★ THE BACKSTOP

If you find a violation of any rule above, email [email protected] with the URL and the issue. We respond within 5 business days and either correct, push back with reasoning, or escalate to a public correction. Every correction we issue is logged at /corrections/ — that page is the canonical source of truth for our track record.

§ 02

How pricing data is collected.

#pricing-data

Pricing for 291 LLM models across 22 providers is verified manually against each vendor's own canonical pricing page. No scrapers. No automated billing-API pulls. No third-party aggregators. The trade-off is honest: slower updates than a scraping pipeline would deliver, but every number has a human who checked the source URL on the date stamped next to the price.

01

Canonical source identified

For each model, the vendor's own canonical pricing page is the source of truth (e.g., anthropic.com/pricing, openai.com/api/pricing). Press releases, blog posts, and third-party aggregators are never used as source — they're often stale or out of date the moment a vendor ships a change.

02

Manual verification

Each price is fetched from the canonical URL via WebFetch, read against the published rate card, and the input rate, output rate, cache rate, and any tier thresholds are recorded by hand into the model's JSON record. One human, one source URL, one verification date — no parser layer between the vendor page and the published number.

03

Diff against snapshot

The new reading is compared against the previous published snapshot. Any change triggers a re-read of the full pricing page (not just the row that moved) to make sure the rest of the rate card didn't change at the same time. Mismatches are reconciled before the new price is published.

04

Publish + log

Approved changes propagate to all dependent pages on the next build. The price-history table records the change with timestamp, old value, new value, and the source URL that was verified. Every price on the site has a verifiable provenance trail. Target latency from a vendor announcement to a published correction is under 24 hours; trigger is a vendor-blog signal or a reader-reported discrepancy.

Source typeExampleRefreshVerification
Vendor pricing pageManual WebFetch read of anthropic.com/pricing, openai.com/api/pricing, etc.On change✓ Manual
Vendor announcementOfficial blog/social signal triggers immediate re-read of the canonical page~24h target✓ Manual
Currency exchangeEvery price is shown in USD. Where a vendor publishes only in its home currency — mostly Chinese labs quoting yuan — we convert at one dated ECB reference rate for the whole catalogue, so rows stay comparable: 6.7209 CNY/USD as of Aug 28, 2026. A vendor that publishes its own USD list is read directly and never converted, because the two are separately set and do not track any exchange rate. Per-country display currency is still Phase 2.Quarterly✓ Manual
Country tax/VATStatic config sourced from official government pages (planned — Phase 2, country pages not yet shipped)~ Planned
Third-party aggregatorsListed as not used as source; only used as a sanity cross-check✗ Not source
§ 03

Where you actually buy.

#access-points

The same model is usually sold through several doors: the vendor's own console, a cloud marketplace, a consumer subscription app, sometimes open weights you can host yourself. Each provider page lists those routes because the door you walk through changes what you pay — and the rate we publish belongs to exactly one of them.

The published per-token rate is always the first-party rate, read from the vendor's own pricing page as described above. Where a route bills differently — a marketplace markup, per-seat licensing, a region premium — the route's card says so rather than silently reusing the direct number.

RouteWhat it isRate on the cardIn the dataset
Direct console / APIThe vendor's own platform — console.anthropic.com, platform.openai.com. The route our published rates describe.First-party✓ Source
Cloud marketplaceBedrock, Azure OpenAI, GCP. Bought under an existing cloud contract; regions and model versions often lag the direct API.Parity or noted~ Noted only
Consumer subscriptionPro / Plus / Max plans. Buys a product with usage limits, not tokens at a rate.Per month✗ Never
Open weights / self-hostLicensed weights you run yourself. No per-token fee; cost becomes your own compute.License✗ Not a rate
Gateways / resellersAggregators and third-party inference hosts. Useful as a route, priced at their own margin.Resale✗ Not source
⚠ A SUBSCRIPTION IS NOT A TOKEN PRICE

A $20/month plan and a $3-per-million-tokens API rate are different units of a different product, and mixing them is the fastest way to publish a number that means nothing. Subscription figures live on the access card and in the subscription vs API calculator, which compares them on a modelled workload. They never enter the per-token dataset, and they never appear in a pricing table as if they were a rate.

Resale prices are never a source. A gateway can name its own margin and change it without the vendor doing anything, so a number read off a reseller tells you about the reseller. We'll name those routes when they matter — MIT-licensed weights served from US/EU regions solve a real data-residency problem — but the price stays theirs, not the vendor's.

Route availability is also dated and specific: which regions a marketplace serves, which model versions it actually carries, whether data stays in-region. Where we can't verify a route's terms first-hand, the card reports what the vendor documents and doesn't dress it up as tested.

§ 04

Where scores come from.

#score-sourcing

We do not run benchmarks ourselves. Every quality score on this site is read off a third-party leaderboard that evaluates models under one standardized harness, and each score carries the leaderboard it came from and the date we read it. We state that plainly because the alternative is worse: a desk this size claiming its own 500-task agent runs would be claiming a compute budget it does not have, and you would have no way to check.

We lean on neutral leaderboards rather than vendor numbers because vendor-published scores are not comparable to each other. Custom scaffolding, tuned prompts and larger context budgets are all normal practice in a launch post, and none of them are labeled. One third-party harness applied to every model is a weaker claim than a self-run — but it is an honest one, and it is the same claim for every model on the board.

SourceWhat we take from itScore cellsStatus
llm-stats.comPer-benchmark neutral leaderboards — coding, knowledge, reasoning, math47✓ Primary
tbench.aiTerminal-Bench v2 — agentic terminal tasks4✓ Primary
Vendor-reportedLast resort when no neutral coverage exists; always carries a visible "vendor-reported" tag0~ Flagged
Self-runReserved in the schema; no self-run score is published today0✗ None

Every cell records its value, source, source URL, the configuration the leaderboard used, a provenance tag, and a verification date. A score whose verification date is older than 30 days is re-pulled before it can be published again; it is never quietly re-served as current. All 51 cells on the board today are neutral third-party results.

⚠ AN ABSENT SCORE IS NOT A LOW SCORE

Missing results are omitted, never estimated, and coverage is computed from what is actually present. A model becomes rankable only once it carries at least three core benchmarks including one coding and one reasoning or knowledge result — below that it can show scores but takes no rank. Closed frontier models and several Chinese vendors are genuinely under-covered on open leaderboards, and filling those gaps with a plausible-looking number would turn a coverage hole into a finding.

§ 05

Benchmark handling.

#benchmarks

We track 8 benchmarks, and they are not all equal. 6 are core: they carry a fixed weight and are the only ones that feed a ranking. The rest are profile-only badges — shown on a model where we have them, never folded into a score. Keeping that line visible matters, because a benchmark quietly promoted from badge to ranking input would change every board on the site without anyone announcing it.

BenchmarkDimensionWeightLeaderboard
LiveCodeBench coding 20% llm-stats ↗
SWE-bench Verified agentic-coding 20% llm-stats ↗
Terminal-Bench v2 agentic 15% tbench.ai ↗
MMLU-Pro knowledge 25% llm-stats ↗
GPQA Diamond reasoning 10% llm-stats ↗
AIME 2025/2026 math 10% llm-stats ↗
SWE-bench Pro agentic-coding profile only scale ↗
Humanity's Last Exam reasoning profile only artificialanalysis ↗

4 of the 6 core tests are flagged as saturated — 2 where the field as a whole has topped out and 2 where only the leaders have converged. We still show them, because a saturated test is useful evidence that a cheap model has caught up with an expensive one. We just don't let a one-point lead on a saturated test decide a ranking, since at that end of the scale the gap between two models is noise rather than signal.

We do not publish a page per benchmark, and there is no reproduction script to link to — see § 04 for why. What each score does carry is the leaderboard it came from, the configuration that leaderboard ran, and the date we read it, printed next to the number wherever it appears.

§ 06

How rankings work.

#rankings

Editorial rankings (Top 10 lists, Best-of picks) follow a strict process: rubric locked → tools scored independently → outliers re-reviewed → final ranking computed. Sponsorship cannot move a tool up or down. Affiliate relationships exist but are disclosed per-link.

01

Rubric locked

For each ranking, we define 6-8 criteria + explicit weights (e.g., Aesthetic 25%, Adherence 20%, Cost 15%). Rubric is published before testing starts. Any change to rubric requires version bump.

02

Pass one — rubric scoring

Founder scores every candidate against the locked rubric, criterion by criterion. Each candidate is scored in isolation; the previous candidate's scores are not visible while scoring the next one. Sponsor status is not visible at this stage. One scorer, one rubric, no shortcuts.

03

Pass two — checklist sweep

A separate checklist pass verifies the scores against the underlying evidence: each criterion must point at a specific source, screenshot, or test run. Anything not anchored to evidence gets flagged for re-scoring. This catches the "I just felt it" kind of mistake.

04

Pass three — cool-off review

After a 48-hour cool-off, the founder re-reads the entire ranking from a fresh head. Score adjustments at this stage are rare but logged with the reason. Final 0-100 score is weighted per rubric and summed; ranking is deterministic from the score. If a sponsor lands in #1 by the math, that's the math. One person, three passes, no manual override.

§ 07

How head-to-heads are built.

#comparisons

A comparison page — there are 48 of them carries no numbers of its own. Each side declares a reference into the pricing snapshot, and every price, context window and effective-cost figure resolves from that snapshot when the site is built. This is deliberate: a comparison that stores its own copy of a price will eventually contradict the model page it links to, and the reader has no way to tell which one is stale.

01

Pair chosen

A pair ships only when both sides can be sourced the same way. Where one side has no neutral coverage — IDE plugins with no agent benchmark that measures both, for example — the pair is held rather than published with one side estimated. Page order follows whichever direction people actually search for; the reverse spelling redirects to it instead of splitting the same content across two URLs.

02

Figures resolve, never typed

Derived rows read the snapshot at build time. If a reference is missing or doesn't resolve, the build fails — it does not fall back to a blank cell or a remembered number. The only hand-written figures are ones the snapshot genuinely can't hold, such as subscription plan prices, and those are marked as authored.

03

Row winners computed

Per-row wins are arithmetic, not opinion: cheaper wins on a rate, larger wins on context. Where one side publishes no cache or batch price at all, the row shows a dash and awards nobody the win — "no published cache rate" is not "infinitely expensive," and treating absence as a loss would quietly punish vendors for thin documentation.

04

Gate before publish

A verification script runs over every comparison before it ships. It blocks unfilled placeholders, references that don't resolve, model pairs with no benchmark section, and out-of-range FAQ counts; it warns on thin editorial depth and on links to pages that don't exist yet. A page that fails the gate does not go out.

Benchmark bars come from third-party leaderboards, joined by the same model reference, with the source and verification date printed under every chart. When a neutral score is missing for either side, the build refuses the card and the author has to cite scores explicitly instead. That rule exists because an absent score is not a low score — closed models are routinely under-covered, and quietly rendering a gap as a short bar would read as a finding when it's a coverage hole.

The verdict table on top is the editorial layer, and it's marked as such. Category calls are a judgement, but each one carries a one-line margin naming the number and the date behind it, so you can disagree with the call while still seeing what it was based on. Where two plans of the same product are compared, benchmarks are dropped entirely: identical underlying models would render as two identical bars and imply a measurement that never happened.

Reader quotes, where a page has them, are real threads only — each with a live link, a date, and a length cap. They're context on how a tool feels in practice, never a substitute for a measurement.

§ 08

Localized pricing verification.

#localized-pricing

We do not publish localized pricing yet. Every rate on the site today is the vendor's list price in USD, shown as the vendor publishes it. No currency conversion, no VAT or GST, no per-country availability — because none of those are things we can currently verify to the standard the rest of this document describes.

This section exists to state the plan and the gate, not to describe something already shipped. When country pages do land, the rule is that a converted price is a derived number and has to be labeled as one: the vendor's USD list price stays visible next to it, with the rate source and the date of conversion. A price that silently becomes €-denominated is a price the reader can no longer check against the vendor's own page.

Tax is the harder half and it is the reason this is gated rather than merely unbuilt. VAT and GST treatment of cross-border API billing depends on the buyer's registration status, not just their country, and getting it wrong produces a number that looks authoritative and is wrong in a way readers can act on. Until that logic is reviewed by someone competent to review it, no tax-inclusive figure ships — an absent country page is a far cheaper mistake than a confidently incorrect tax rate.

§ 09

How calculators work.

#calculators

Every calculator uses verifiable math — no black-box estimation. The formula is printed on each calculator page, every rate comes from the same vendor-verified pricing snapshot as the rest of the site, and all numbers on a calculator page (presets, worked examples, comparison tables) are recomputed from that snapshot on every rebuild — nothing is hand-typed.

The LLM API cost calculator models the levers that actually move bills: prompt-cache hit rates with vendor cache-write premiums, batch-tier rates, and hidden reasoning tokens billed at the output rate. Default assumptions are shown in the UI, not hidden, and every input can be tweaked.

Token counting is labeled with one of three accuracy tiers, on every row. exact — OpenAI models counted in your browser with the o200k_base BPE, the same encoding the API uses. cal — vendors whose token counts use a chars-per-token calibration we measured on the vendor's own published tokenizer (downloaded from their official Hugging Face repos; measurement date shown on the token counter). est — vendors with no public tokenizer (Claude, Gemini, Grok, Kimi), estimated from vendor documentation. Where a closed API model is calibrated via the vendor's open-weight sibling, the page says so.

When vendors reprice or release new tokenizers, the snapshot and calibrations are re-verified and the whole section recomputes on deploy. Found math that looks wrong? Report it — calculator errors are treated as corrections, not feedback.

§ 10

Capability tests.

#capability-tests

For "Can X do Y?" pages (e.g., "Can Claude Code work offline?"), we run the test ourselves, capture the actual output, and label answers with verification status. The tested version, date, and exact command/prompt used are shown on every capability page.

Where the answer is "no," we always provide workarounds — the alternative path or tool that does support the capability. We don't just confirm a limitation and walk away. The capability question template's "Alternative path" section is required for any "no" answer.

§ 11

Install guides.

#install-guides

For every install guide, the author installs the tool on a clean machine on each platform we cover (macOS, Linux, Windows). Steps are recorded as the install runs; expected output is captured from actual terminal output, not described from memory. Tested-with version is logged in the byline.

Guides are re-verified quarterly or whenever a tool ships a new major version. Common errors come from our own install logs plus support-channel monitoring on each tool's Discord/forum/Stack Overflow tag — we add troubleshooting items as we encounter or hear about them.

§ 12

Sponsorship policy.

#sponsorship

Sponsored slots exist on some pages. They are visually quarantined (warm-amber dashed border, "ⓘ SPONSORED" tag, never inside editorial ranking blocks) and cannot affect editorial scoring. Sponsorship is sold at fixed published rates.

Affiliate relationships are separate from sponsorship. Many tools we recommend have affiliate programs we participate in. Each affiliate link is marked with the ↗ symbol. Affiliate revenue funds our testing budget. Affiliate status cannot move a tool up or down in editorial rankings — and we cite tools without affiliate programs (DALL-E, Imagen, OpenAI direct API) on their merits when they win on rubric.

⚠ WHAT SPONSORS CANNOT BUY

Top-N positions, "best of" picks, scoreboard rankings, brand profile verdict, comparison verdicts. Sponsors can buy quarantined slots in clearly-labeled sponsored sections. They cannot buy editorial position. The policy is non-negotiable, and it is written down here before the first slot is sold rather than after.

§ 13

How corrections work.

#corrections

When we get something wrong, we log it. The corrections page lists every correction with: date issued, page affected, what was wrong, what's now correct, who flagged it. We don't silently edit. We don't claim things were always correct that were not.

Severity levels: minor (typo, broken link) → fixed silently with timestamp on page. Material (incorrect price, wrong benchmark score, wrong feature claim) → corrected page + entry in /corrections/ + email reply to whoever flagged it. Severe (incorrect ranking, retracted claim) → corrected page + corrections log + apology paragraph in next quarterly editorial review.

The canonical record of every correction we've issued lives at /corrections/ — each entry includes the original claim, the corrected claim, the date, and the page affected.

§ 14

The desk.

#reviewers

Founder signs off on every published claim. AI assists with drafting and formatting; the human author makes every editorial decision and verifies every fact. As the desk scales, additional named reviewers will appear here with verifiable handles — second-pair-of-eyes hires are welcome.

Y

Yaroslav VikharievFounder

Founder of AI Cost & Tools Hub. Editorial sign-off on every pricing claim. Verifies prices against the vendor's own canonical pricing page; leaves fields blank rather than fabricating.

§ 15

Version log.

#changelog

Every change to this methodology is logged. The page you're reading is v1.3. The log below starts where the document does — this site published its first page in May 2026, and there is no earlier methodology to point at.

v1.3
JUL 20 · 2026
Added § 07 how head-to-heads are built · Rewrote § 04 and § 05 to describe where benchmark scores actually come from — the previous text claimed we ran SWE-bench, GAIA and WebArena ourselves, which was never true · Removed § 08 claims about shipped country pages and named accountant review.
v1.2
JUN 10 · 2026
Added § 09 token-counting accuracy tiers (exact / calibrated / estimated) · Updated § 02 to describe manual per-vendor verification rather than an automated parser layer.
v1.1
JUN 01 · 2026
Removed unverified team and service claims carried over from the page mockup · Changed model and provider counts to derive from the catalog itself so the prose cannot drift.
v1.0
MAY 17 · 2026
First public methodology, published with the site.