Aggregated from 10 of our own comparisons
Best AI coding agents.
108 judged categories · 10 head-to-heads · no score of ours
This table is arithmetic on comparisons we already published, not an opinion poll. Ten agent pairs, 108 categories, each judged against vendor pages and neutral benchmarks on a dated page you can open. Two things are wrong with the method and both are stated below. The tool we cover in most depth does not come first.
Ranked on categories won.
Only tools with 2 or more published head-to-heads. The opponents column is the point: a rate is only as meaningful as who produced it.
| # | Tool | Categories won | Rate | Measured against |
|---|---|---|---|---|
| 01 | GitHub Copilot | 15 of 24 · 1 tied | 63% | Claude Code·Cursor |
| 02 | Cursor | 28 of 67 · 5 tied | 42% | Claude Code·Cline·Codex·GitHub Copilot·Windsurf·VS Code |
| 03 | Codex | 7 of 21 · 6 tied | 33% | Claude Code·Cursor |
| 04 | Claude Code | 12 of 39 · 7 tied | 31% | Cursor·Codex·GitHub Copilot·Gemini CLI |
Measured once, so not placed.
These have a published comparison but only one, and one pair cannot separate the tool from its opponent. They are here so the list is not silently shorter than the evidence.
| Tool | Categories won | Measured against |
|---|---|---|
| Cline | 8 of 12 | Cursor |
| Gemini CLI | 4 of 8 | Claude Code |
| Lovable | 4 of 11 | Replit |
| Replit | 6 of 11 | Lovable |
| VS Code | 6 of 11 | Cursor |
| Windsurf | 6 of 12 | Cursor |
It is not a score we assigned
Every row is a count of category verdicts from comparisons that already exist on this site, each with its own verification date. We publish no composite rating and run no benchmark of our own — where a comparison cites a benchmark, it cites somebody else's and says whose.
The schedule is not balanced
Tools were not all measured against the same opponents, so a rate carries the strength of whoever it was compared with. That is why the minimum exists and why the opponents are printed next to every rate rather than summarised away.
Categories are not weighted
Cost per task counts the same as code quality, and there is no finer view to fall back on: the ten pairs use 89 distinct category names across 118 rows, so nothing aggregates cleanly below this level. The pair pages are the real resolution.
Nobody is crowned
Every one of our comparisons ends in "it depends" and a decision tree, on purpose — which tool wins depends on what you are doing. That is why this page counts categories instead of victories: there are no victories to count.
It is not for sale
No entry paid to be here and we sell none of these tools. The order is recomputed from the underlying comparisons on every build, so it changes when the evidence does and not when someone asks.
Frequently asked.
What people ask about a ranking like this.