R’s Perspective
All writing

Sonnet 5.5 vs 5

Four runs each. The same fix.
A much shorter wait.

2 min read

I ran four rounds of tests with Sonnet 5.5 and Sonnet 5, asking each to fix the same issue in an individual worktree. Sonnet 5.5 finished every run sooner.

Across the four runs, the average execution time was 4.8 minutes for Sonnet 5.5 and 15.9 minutes for Sonnet 5. That is roughly 3.3 times as fast, with about half as many turns.

The task

The view renderer and the copilot each had their own checks for draft access, required permissions, and pinned-view ownership. Keeping two copies had already led to a missed permission check in the copilot.

Each agent needed to extract one shared function and test both callers. The renderer’s 404 response, the copilot’s empty query set, and the existing trace output all had to stay intact.

Same commit, same brief, same repo. Each model ran the ticket four times in its own worktree, using a headless Claude Code programmer at high effort. A third model reviewed every diff blind.

The results

The shorter run time came with fewer turns, fewer tool calls, and lower token use and cost. These are the rounded averages across the four runs.

Average across four runs per model (29 September 2026)
MetricSonnet 5.5Sonnet 5
Execution time4.8 min15.9 min
Turns51106
Tool calls64131
Tokens23.6k63.3k
Cost$0.93$3.36
Sonnet 5.5 and Sonnet 5 benchmark results for four runs of the same refactor, showing execution time, turns, tool calls, tokens, cost, and a per-run execution-time chart.
The full four-run comparison. Click the image to view it at full size.

Time, cost, and blind review

The median tells a similar story to the averages: 4.6 minutes per run for Sonnet 5.5 versus 15.3 minutes for Sonnet 5. Median cost was $0.88 versus $3.33. On this ticket, that made Sonnet 5.5 about 3.3 times as fast and 3.8 times as cheap.

The independent blind review also favoured Sonnet 5.5. Its mean score was 17.3 out of 20, compared with 14.8 for Sonnet 5 — a 2.5-point difference. The faster runs came with a higher review score, not just a shorter wait.

What the comparison tells me

This is a small comparison on one concrete issue, with four runs per model. The timing, cost, and blind-review results all favoured Sonnet 5.5 here. That gives me a useful signal for this kind of refactor, rather than a result for every kind of coding task.