Last verified
CompareObjective verdictsLive pricing

Compare48

Side-by-side AI model head-to-heads — live pricing, neutral benchmarks, and a clear verdict by use case. We call the winner where there is one, and say "tie" where there isn't.

Head-to-heads
48
2-way pairs
Models compared
13
across all pairs
Providers
19
brands covered
Benchmarks
6core
neutral leaderboards
Pricing
live
from vendor pages
Verified
08/31
2026
§ 02 / HEAD-TO-HEADS

Model head-to-heads.

2-WAY
Antigravity vs Cursor

Antigravity is free and generally available; Cursor starts at $20. On Artificial Analysis's Coding Agent Index, Google's harness scores 56.6 against C…

Google vs Anysphere
2-WAY
Claude Code vs GitHub Copilot

Claude Code is a terminal agent running Anthropic models, and it tops the neutral coding-agent leaderboards. GitHub Copilot is a layer across your edi…

Anthropic vs GitHub
2-WAY
Claude Sonnet vs Claude Haiku

Same family, two different jobs. Claude Haiku is half the price ($1 / $5 vs $2 / $10 per 1M) and the fastest tier on Anthropic's own latency table. Cl…

Anthropic vs Anthropic
2-WAY
Claude Sonnet vs Claude Opus

Same family, two jobs. Claude Sonnet is the 2.5× cheaper workhorse ($2 / $10 vs $5 / $25 per 1M) that handles roughly 90% of coding and chat work. Cla…

Anthropic vs Anthropic
2-WAY
Claude vs ChatGPT

Both are frontier assistants that trade wins. Claude leads real-world coding (SWE-bench Pro 69.2 vs 63.4), long-form writing, and is MCP-native. ChatG…

Anthropic vs OpenAI
2-WAY
Claude vs Gemini

These two split cleanly. Claude Opus 4.8 wins quality — real-world coding (SWE-bench Pro 69.2 vs 55.1), frontier reasoning, long-form writing. Gemini …

Anthropic vs Google
2-WAY
Cline vs Cursor

Two cost models, not two intelligences. Cline is a free Apache-2.0 extension on bring-your-own-key: you pay the provider directly, on any model, with …

Cline Bot vs Anysphere
2-WAY
Codex vs Claude Code

Both are terminal-plus-cloud coding agents that run neck-and-neck. Codex (OpenAI, GPT-5.x-Codex) has more surfaces — CLI, IDE, cloud, GitHub, iOS, Chr…

OpenAI vs Anthropic
2-WAY
Codex vs Cursor

Codex is the stronger agent and Cursor is the cheaper one, by wide margins in both directions. On the Artificial Analysis Coding Agent Index, Codex on…

OpenAI vs Anysphere
2-WAY
Microsoft Copilot vs ChatGPT

This is Microsoft's consumer assistant, not GitHub Copilot. Copilot is free, built into Windows, and its paid path — Microsoft 365 Premium at $19.99/m…

Microsoft vs OpenAI
2-WAY
Copilot vs Gemini

Both sell AI as an entitlement on a productivity subscription, and both charge $19.99 for the main tier. Gemini names its models, ships a real free ti…

Microsoft vs Google
2-WAY
Cursor vs Claude Code

Two different shapes of tool. Cursor is an AI-native IDE that runs Claude, GPT, Gemini or its own Composer; Claude Code is a terminal-native agent on …

Anysphere vs Anthropic
2-WAY
Cursor vs GitHub Copilot

This is GitHub Copilot, the coding tool — not Microsoft's consumer assistant. GitHub Copilot is half the entry price at $10/mo, $19 a team seat, and w…

Anysphere vs GitHub
2-WAY
Cursor vs VS Code

Cursor is a fork of VS Code, so the editor underneath is the same. What you pay for is the AI layer: a house model at $0.55 per agent task, per-hunk m…

Anysphere vs Microsoft
2-WAY
DeepSeek vs ChatGPT

DeepSeek and ChatGPT answer different questions. DeepSeek V4 Pro is open-weight and radically cheaper — the API runs ~6× cheaper on input and ~17× che…

DeepSeek vs OpenAI
2-WAY
DeepSeek vs Claude

DeepSeek and Claude sit at opposite ends of the price/quality curve. Claude Opus 4.8 wins quality — GPQA Diamond 93.6 vs 90.1, SWE-bench Verified 88.6…

DeepSeek vs Anthropic
2-WAY
DeepSeek vs Gemini

DeepSeek sells price and control; Gemini sells reach. DeepSeek V4 Pro is open-weight (MIT) and ~3.4× cheaper on input, ~9× cheaper on output, and you …

DeepSeek vs Google
2-WAY
Gemini CLI vs Claude Code

Gemini CLI is free and open-source — Apache-2.0, roughly 1,000 requests/day on a Google account, backed by Gemini 3.1 Pro's 1M-token, multimodal conte…

Google vs Anthropic
2-WAY
Gemini Pro vs Gemini Flash

Google's tier ladder inverted. Gemini 3.6 Flash (GA, Jul 21) is cheaper ($0.75 / $3.75 vs $2 / $12), about 2× faster, and outscores the older Gemini 3…

Google vs Google
2-WAY
Gemini vs ChatGPT

Gemini and ChatGPT are the two default assistants, and they split the field. Gemini is cheaper on the API ($0.75 / $3.75 vs $2.50 / $15 per 1M), has a…

Google vs OpenAI
2-WAY
Grok vs ChatGPT

Both are frontier assistants that trade wins. Grok is cheaper on the API ($2 / $6 vs $2.50 / $15 per 1M tokens), native to real-time X, and does video…

SpaceXAI vs OpenAI
2-WAY
Grok vs Claude

On raw reasoning they are level — GPQA Diamond 93.0 vs 93.6, Terminal-Bench 2.1 79.3 vs 78.9, both inside the error bars. They split on economics and …

SpaceXAI vs Anthropic
2-WAY
Grok vs Gemini

Grok and Gemini split the field. Grok owns real-time X and leads Artificial Analysis' composite index (54 vs 50). Gemini has double the context (1M vs…

SpaceXAI vs Google
2-WAY
Kimi vs Claude

Kimi K2.6 is open-weight and roughly 5× cheaper on input, 6× on output than Claude Opus 4.8. Claude wins quality — SWE-bench Verified 88.6 vs 80.2, GP…

Moonshot AI vs Anthropic
2-WAY
Kimi vs Qwen

Two Chinese labs at the frontier, priced in opposite directions. Qwen3.8 Max lists $2/$6 against Kimi K3's $3/$15, so raw output is 2.5× cheaper. But …

Moonshot AI vs Alibaba
2-WAY
LM Studio vs Ollama

Both run open-weight models on your own hardware, and on one machine both can load the same GGUF file. LM Studio is a desktop app with a model browser…

LM Studio vs Ollama
2-WAY
Ollama vs llama.cpp

This is not two competing runtimes. Ollama's own README lists llama.cpp under "Supported backends", so the question is what the wrapper adds. Less tha…

Ollama vs ggml-org
2-WAY
Perplexity vs ChatGPT

Two different tools, not two models. Perplexity is an answer engine: every response is web-grounded with clickable citations by default, and Pro route…

Perplexity vs OpenAI
2-WAY
Perplexity vs Claude

Perplexity and Claude solve different problems. Perplexity is a cited answer engine — real-time web search, inline sources, and Deep Research at $20/m…

Perplexity vs Anthropic
2-WAY
Perplexity vs Gemini

An answer engine against a general assistant. Perplexity grounds every reply in live sources with inline citations by default, and Pro routes your que…

Perplexity vs Google
2-WAY
Qwen vs ChatGPT

Qwen and ChatGPT split the board. Qwen Studio is free with no consumer paid tier, and Qwen3.7 Max is cheaper on output tokens. ChatGPT leads on compos…

Alibaba vs OpenAI
2-WAY
Qwen vs Claude

Qwen and Claude both run 1M-context flagships, and they split cleanly. Claude Opus 4.8 wins quality: SWE-bench Verified 88.6 vs 80.4, GPQA Diamond 93.…

Alibaba vs Anthropic
2-WAY
Qwen vs DeepSeek

China's two biggest model families, and they diverge sharply. Qwen 3.7 Max leads reasoning and knowledge — GPQA 92.4 vs 90.1, MMLU-Pro 89.6 vs 87.5 — …

Alibaba vs DeepSeek
2-WAY
Replit vs Lovable

Both turn a prompt into a running app and bill by credits, but the meters differ. Replit denominates credits in dollars — $25 of credits on a $25 plan…

Replit vs Lovable
2-WAY
vLLM vs Ollama

These are not rivals so much as different jobs. vLLM is a serving engine: continuous batching, PagedAttention, tensor and pipeline parallelism, Linux …

vLLM project vs Ollama
2-WAY
Windsurf vs Cursor

Two VS Code forks that run the same frontier models. Windsurf now ships as Devin Desktop under Cognition, and it wins on cost mechanics — unlimited Ta…

Cognition vs Anysphere
§ 02 / HEAD-TO-HEADS

Plan & tier comparisons.

PLAN VS PLAN
ChatGPT Free vs ChatGPT Paid

ChatGPT Free runs GPT-5.5 Instant with capped messages, may show ads in the US, and marks most features 'limited' — no GPT-5.6 reasoning models and no…

OpenAI vs OpenAI
PLAN VS PLAN
ChatGPT Go vs ChatGPT Plus

ChatGPT Go ($8/mo) is Free with roughly 10× the headroom — more messages, uploads, and image generation — but it keeps GPT-5.5 Instant only, stays ad-…

OpenAI vs OpenAI
PLAN VS PLAN
ChatGPT Plus vs ChatGPT Pro

ChatGPT Plus ($20/mo) and Pro ($200/mo) run the same models, but Pro unlocks the deepest reasoning — GPT-5.6 Sol Pro plus Extra High — 20× the usage, …

OpenAI vs OpenAI
PLAN VS PLAN
ChatGPT Plus vs ChatGPT Team

ChatGPT Plus is $20/mo for one person. ChatGPT Team — renamed ChatGPT Business — is $20/user/mo billed annually ($25 monthly, 2-seat minimum). Same 54…

OpenAI vs OpenAI
PLAN VS PLAN
Claude Free vs Claude Pro

Since June 30, 2026 Free and Pro run the same default model, Sonnet 5 — so raw answer quality is identical. Pro's $20/mo buys roughly 5× the usage, Op…

Anthropic vs Anthropic
PLAN VS PLAN
Claude Max 5x vs Max 20x

Two rungs of one plan. Max 5x is $100 a month, Max 20x is $200 — and across all 53 rows of Anthropic's own plan-comparison table, every other cell is …

Anthropic vs Anthropic
PLAN VS PLAN
Claude Pro vs Claude Max

Same Claude, same Opus 4.8 — the only difference is how much you can run. Pro is $20/mo and enough for most; Max is $100 (5×) or $200 (20×) the per-se…

Anthropic vs Anthropic
PLAN VS PLAN
Claude Team vs Claude Enterprise

Team is a listed price with usage baked in: $20–$25 a standard seat, $100–$125 premium, minimum two members, and heavy weeks get throttled. Enterprise…

Anthropic vs Anthropic
PLAN VS PLAN
Gemini Free vs Google AI Pro

Same models, different headroom. Gemini Free already runs Gemini 3 Pro at $0, but caps you at a 32k context window and standard usage. Google AI Pro (…

Google vs Google
PLAN VS PLAN
Google AI Pro vs Google AI Ultra

Same Gemini models, different headroom. Google AI Pro is $19.99/mo — 4x free limits, 1,000 Flow credits, 5 TB. Google AI Ultra starts at $99.99/mo (5x…

Google vs Google
PLAN VS PLAN
Grok Free vs SuperGrok

Grok Free ($0) gives you Grok, voice, connectors and limited real-time X search. SuperGrok ($30/mo) unlocks the Grok 4.5 frontier model, Expert mode, …

SpaceXAI vs SpaceXAI
PLAN VS PLAN
Perplexity Free vs Perplexity Pro

Perplexity Free and Pro run the same engine — the gap is volume and access. Free caps you at 3 Pro Searches a day, one Research report a month, and no…

Perplexity vs Perplexity

One weekly digest. Zero noise.

One weekly digest · No spam, ever · Unsubscribe in one click