ai coding agents

Best Model for Coding: What 1,232 Reviewed PRs Show

By Ilya Volodarsky | | 8 min read

Related guide: Claude Code vs Codex for teams

Our team writes code with Claude Opus 5.5 in Claude Code, and has since the day it shipped, September 22, 2026. It costs less per token than Fable 5.1, which it replaced, and its pull requests merged fastest. But the bigger lesson from our own data is that the model matters less than you’d think: once we compared the same engineer, repository and week, the gap between models mostly disappeared, and the size of the pull request predicted review findings far better.

This post is the model side of what 866 agent-written pull requests show. All our numbers come from two repositories at Winding Labs, merged from August 12 to October 4, 2026.

What we use, and why

Our default changed twice in eight weeks:

  • August: Fable 5 and Opus 5, side by side.
  • From September 1: Fable 5.1.
  • From September 22: Opus 5.5. Every one of the 74 pull requests merged on October 1–4 ran on it.

Codex wrote 31 of our 866 agent PRs, all in September. Claude Code wrote 822.

Three things keep us on Opus 5.5:

  • Price per token. Opus 5.5 lists at $4 per million input tokens and $20 per million output tokens. Fable 5.1 lists at $10 and $50. Cache reads cost $0.20 and $0.25 per million, and they are most of an agent’s tokens: 98% of one engineer’s in September.
  • Price per merged PR. On one engineer’s laptop, from September 4 to October 3, a merged PR cost $19 in Opus 5.5 tokens at list price (374 PRs) and $27 in Fable 5.1 tokens (224 PRs). Opus 5.5 used more tokens per PR, 68 million against 47 million, and still cost less.
  • Speed to merge. The median Opus 5.5 PR merged 25 minutes after it was opened, against 41 for Fable 5.1, on PRs of about the same size.

Cheaper tokens did not make the month cheaper. In the same engineer’s weeks, the list price per million tokens fell from $0.59 to $0.30 after the switch, while weekly use rose from 7 to 9 billion tokens to 30 billion. The weeks held different work, so this is a correlation, not a measured effect. It is still the pattern to watch: a cheaper model gets used more.

Our numbers by model

Each row counts PRs whose agent named that model in the block it writes on every pull request. The last column is the share our AI reviewer, IonWarp (our sister product), flagged with a should-fix issue on its first pass.

Model PRs Median time to merge Median lines changed Flagged on first review
Opus 5.5 213 25 min 693 32% of 214
Fable 5.1 255 41 min 697 31% of 169
Fable 5 176 33 min 611 48% of 153
Opus 5 149 47 min 580 55% of 127
Codex 31 37 min 576 23% of 22

The review counts come from a later pull of the same window and cover the PRs with a readable first pass, so they differ from the PR counts. Most Codex blocks don’t name the model.

Read the last column with care. Each model dominated different weeks, and the reviewer changed week to week too: its format, its models and how strict it was. A model that ran in a strict week looks worse for reasons that have nothing to do with the model.

Matched, the model gap mostly disappears

So we compared models only inside the same repository, engineer and week, and took a weighted difference with a bootstrap 95% confidence interval:

Comparison (same repo, engineer, week) PRs Difference in first-review flag rate 95% interval
Codex vs Claude Code 24 vs 162 +1 point −16 to +20
Opus 5.5 vs Fable 5.1 (week of Sep 21) 100 vs 20 −7 points −30 to +15
Fable 5 vs Opus 5 (August) 152 vs 103 −12 points −26 to 0
PRs of 1,000+ lines vs under 300 234 vs 242 +36 points +27 to +43

Only one model gap survives, and only just: in August, Fable 5 PRs were flagged a little less often than Opus 5 PRs. Codex against Claude Code is a coin flip at this sample size. Size is the one row that is large and certain.

Across all 1,232 reviewed PRs, the reviewer flagged 11% of PRs under 100 changed lines and 57% of PRs between 1,000 and 3,000. The gradient holds in August alone and in late September alone. The full table is in the 866-PR post.

The practical reading: switching models will move your review findings by a few points at most, while asking the agent for smaller PRs moves them by tens of points.

What public benchmarks say

Public benchmarks answer a narrower question than “which model writes the best code for us”, but they are the right place to build a shortlist.

Artificial Analysis Coding Agent Index v1.5 runs coding agents on three benchmarks (DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA). The entries highlighted on its chart, read on October 5, 2026:

Agent and model Index Average cost per task
Claude Code, Sonnet 5.5 (max) 68.4 $14.19
Claude Code, Opus 5.5 (max) 66.0 $13.04
Codex, GPT-6.1 Sol (xhigh) 62.9 $1.04
Claude Code, Fable 5.1 (max) 62.2 $12.39
Codex, GPT-6 Astra (max) 61.6 $7.47
Grok Build, Grok 4.7 (xhigh) 56.3 $8.82
Muse Code, Muse Spark 1.3 (max) 54.3 $3.98

Two things stand out. The top entries sit within six points of each other while their cost per task differs more than tenfold. And effort moves the score as much as the model does: on the same page, Claude Code with Sonnet 5.5 scored 68.4 at max effort and 62.9 at xhigh, for $3.33 a task. Sonnet 5.5 came out on September 28, 2026, and we haven’t run it on our own repositories yet, so we have no data of our own on it.

SWE-bench Verified, the benchmark most articles still quote, had no entry newer than February 26, 2026 when we read it. None of the models above is on it, so a ranking built from it is a ranking of last winter’s models.

Reviewing is a different question. Our sister product publishes a code-review benchmark, marked provisional (snapshot captured September 25, 2026): models review 50 open-source pull requests with 158 known issues. The best model found 79 of them with 42 false alarms, at $0.42 a review. That ranks models as reviewers, not as authors, so it doesn’t answer which model should write your code.

How to pick a model for your team

  1. Shortlist two from a current benchmark. Take one top model and one that costs a fraction per task. Check the date on the leaderboard before you trust it.
  2. Run both on the same work, in the same weeks. Our raw numbers said one model was flagged twice as often as another; matched by engineer and week, the gap was a few points. If one model runs in August and the other in September, you are measuring the calendar.
  3. Record the model on every session. Ours changed twice in four weeks. A monthly “per model” number goes stale fast unless each session stores its model.
  4. Judge per merged PR, not per token. Cost per merged PR, time to merge, review findings and fix rounds. A cheaper token that needs more turns is not cheaper.
  5. Keep pull requests small, whatever the model. It moved our review findings more than any model switch did.
  6. Check again next month. Anthropic released six Claude models between June 9 and September 28, 2026, from Fable 5 to Sonnet 5.5.

Zest records the model behind every Claude Code and Codex session and joins each session to the pull request it led to, so you can run this comparison on your own repositories: see Claude Code analytics and Claude Code vs Codex for teams.

Caveats

  • One small team, two repositories. Most dash PRs come from one engineer. This is how we work, not a sample of the industry.
  • Different weeks, different work. The matched comparisons fix that, but their samples are small and their intervals wide.
  • Agent-written labels. The harness and model come from a block the agent writes itself.
  • Our own reviewer. IonWarp is our sister product, and it changed during the window. 290 of 1,569 reviewed PRs had no readable first pass.
  • List price, one engineer. Cost per PR is what the tokens would cost on the API, for one heavy user on a subscription, not what we paid.
  • Benchmarks move weekly. The public numbers above were read on October 5, 2026.

Frequently asked questions

Which Claude model is best for coding?

On Artificial Analysis’ Coding Agent Index, read on October 5, 2026, Claude Sonnet 5.5 at max effort scored highest of the Claude models in Claude Code (68.4), ahead of Opus 5.5 (66.0) and Fable 5.1 (62.2). On our own repositories we use Opus 5.5: it costs less per token than Fable 5.1 and its pull requests merged fastest. We have no data of our own on Sonnet 5.5 yet.

What is the best LLM for coding right now?

The top of the public coding-agent benchmarks is close. On October 5, 2026, Claude Code with Sonnet 5.5 or Opus 5.5 and Codex with GPT-6.1 Sol were within six points of each other on Artificial Analysis’ Coding Agent Index, while their average cost per task ranged from about $1 to $14. Shortlist two and measure them on your own pull requests.

Does the model matter more than pull request size?

Not on our repositories. Matched on engineer, repository and week, models differed in first-review flag rate by a few points with wide intervals, while pull requests of 1,000 lines or more were flagged 36 points more often than pull requests under 300 lines.

Is Codex better than Claude Code for writing code?

On our repositories we couldn’t tell them apart. Matched on engineer, repository and week, Codex pull requests were flagged on first review at almost the same rate as Claude Code pull requests: +1 point, with an interval of −16 to +20, from only 24 Codex pull requests.

Sources