
Sonnet 5.5 vs Opus 5.5: which slot each one takes in a coding agent stack
Tools in this post
Claude Sonnet 5.5, released on 28 September 2026, costs $2 and $10 per million input and output tokens. Claude Opus 5.5, released six days earlier, costs $4 and $20. Cache reads cost $0.20 per million on both, so inside a coding agent the gap is far smaller than 2x: one metered Sonnet 5.5 run that priced out at $35.40 would have cost $50.48 at Opus 5.5 rates, a 30% saving. At medium effort, the Claude Code default, Anthropic's own chart has Sonnet 5.5 at 28.8% on Terminal-Bench 4.0 against 57.6% for Opus 5.5.
Those numbers point to a split: Opus 5.5 in the planner and reviewer seat, Sonnet 5.5 writing edits that are already scoped. This post covers the seats, as a follow-up to the model routing post and the Opus 5.5 harness checklist.
What is the difference between Claude Sonnet 5.5 and Claude Opus 5.5?
Price, speed and how much each one does at low effort. Both have a 1M-token context window, 128K of output and a June 2026 knowledge cutoff. Sonnet 5.5 costs half as much per token on everything except cache reads, and generates output faster. Opus 5.5 scores far higher at low and medium effort. Near the top of the effort scale the two converge, in score and in cost.
| Sonnet 5.5 | Opus 5.5 | |
|---|---|---|
| Released | 28 September 2026 | 22 September 2026 |
| API ID | claude-sonnet-5-5 |
claude-opus-5-5 |
| Input / output, per million tokens | $2 / $10 | $4 / $20 |
| Cache read | $0.20 | $0.20 |
| Cache write, 5 minutes / 1 hour | $2.50 / $4 | $5 / $8 |
| Default effort on the API | high |
medium |
| Default effort in Claude Code and the apps | medium |
medium |
| Fast mode | No | Yes, at $8 / $40 |
The cache read row is the odd one. Anthropic's pricing page says "Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price", where every other model pays 0.1x. Both come to $0.20.
On speed, Artificial Analysis measured 92.0 output tokens per second for Sonnet 5.5 at medium effort against 73.2 for Opus 5.5. The API defaults also differ, so a "default against default" API test runs Sonnet at high against Opus at medium.
Which of the two should plan and which should write the edits?
Opus 5.5 plans and reviews, and Sonnet 5.5 writes the edits once they are scoped. That is Anthropic's own positioning, and the two independent tests published so far point the same way. Nobody, Anthropic included, has published a measurement of an Opus 5.5 planner with a Sonnet 5.5 executor, so the split is a reasonable default and not a proven one.
Anthropic's launch page says: "Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks". It adds that Sonnet 5.5 "complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost."
| Test | Sonnet 5.5 | Opus 5.5 | Measured by |
|---|---|---|---|
| Code review, 13 hard known bugs | 6 caught, 41.2% precision | 8 caught at standard with 66.7% precision, 10 at max | CodeRabbit, which sells review built on these models (post, 28 September). The Opus figures are from its earlier run on the same cases |
| Implementer agents, one night | $0.68 per 100 merged lines, 13% passed first review | $1.20, 33% passed | Anton Shubin, 48 and 33 pull requests, 30 September |
CodeRabbit concludes that Sonnet 5.5 is "not an Opus 5.5 replacement", because four of its 13 cases defeated every Sonnet configuration. Shubin moved his implementers to Sonnet 5.5 and kept Opus 5.5 for "every review, UI component work and security-sensitive work". He calls it one night of data and plans to rerun it.
Claude Code pairs them in four ways, per its model and advisor docs: opusplan runs Opus in plan mode and Sonnet in execution, /advisor opus lets a Sonnet session consult Opus at decision points, a subagent takes model: sonnet, and /model switches by hand.
Switching has a cost. The caching docs say "each plan-mode toggle is a model switch and starts a fresh cache", so the next request reads the whole conversation with no cache hits. Sonnet 5.5 cannot read Opus 5.5's thinking blocks, so the executor works from the written plan and not from the reasoning behind it. On Amazon Bedrock and Google Cloud the sonnet alias still resolves to Sonnet 4.5, so opusplan there pairs Opus 5.5 with an older model unless you set ANTHROPIC_DEFAULT_SONNET_MODEL.
Anthropic's cost guide, which does not cover Sonnet 5.5 yet, gives the condition: "An orchestrator buys something only when there is bulk to hand off". For one dependent chain of work that fits in a single context, its measurements favoured one model at lower effort. The same two-model switch exists in Aider's architect mode, Cline's Plan and Act modes and OpenCode's per-agent models.
How much do you save switching to Sonnet 5.5 when most tokens are cache reads?
About 30% on identical tokens when cache reads are 57% of the bill, and 20% when they are three quarters of it. Every token class costs twice as much on Opus 5.5 except cache reads. But the two models do not spend the same tokens on the same task, so measured results run from 56% cheaper at medium effort to 28% more expensive at max.
The identical-tokens case comes from a Reddit post of 29 September 2026. One Sonnet 5.5 session in Claude Code made 779 tool calls over about 2.4 hours on a Max 20x plan. The author priced the transcript at API list prices, then priced the same tokens at Opus 5.5 rates. Only the Sonnet run was metered. Recomputed from the pricing page:
| Token class | Tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Output | 739k | $7.39 | $14.78 |
| Cache reads | 101.6M | $20.32 | $20.32 |
| Cache writes, 5 minute | 3.07M | $7.68 | $15.35 |
| Total | $35.39 | $50.45 |
The post says $35.40 and $50.48, so its totals reproduce to within three cents. Its "about 43%" is the Opus premium: $50.48 is 43% more than $35.40. The saving from switching to Sonnet is 30%, not the "half price" in the post's title.
The rule behind it: if cache reads are share s of the Sonnet bill, the same tokens cost 2 minus s times as much on Opus 5.5. At 57% that is 1.43 times. At 75%, the share Shubin reports for his implementers, it is 1.25 times, a 20% saving. With no caching it is 2 times.
Two caveats. The table has no uncached input and assumes 5-minute cache writes. With 1-hour writes throughout, which Claude Code uses for the main conversation on a subscription, the saving would be 33%. And the author says in a comment that the same video on Opus 5.5 "took ~5% of weekly limit and sonnet 5.5 around 2%", a bigger gap than his dollar table shows.
When each model spends its own tokens:
| Measurement | Sonnet 5.5 | Opus 5.5 | Sonnet against Opus |
|---|---|---|---|
| Artificial Analysis, cost per index task, medium | $0.59, index 41 | $1.34, index 51 | 56% cheaper, 10 points lower |
| Artificial Analysis, cost per index task, max | $7.67, index 56 | $5.98, index 58 | 28% more expensive |
| Shubin, per 100 merged lines | $0.68 | $1.20 | 43% cheaper, fewer first-review passes |
At max effort Sonnet 5.5 wrote about 60% more output tokens per task than Opus 5.5 in Artificial Analysis's runs. The saving holds at low and medium effort and disappears at the top.
What did Anthropic measure on Terminal-Bench 4.0 and GDPval-AA?
On Terminal-Bench 4.0, Anthropic reports 70.6% for Sonnet 5.5 at max effort and 66.4% for Opus 5.5 at xhigh, from 66 tasks run five times each in Claude Code. The standard errors are 2.5 and 2.6 points, so the 4.2-point lead is about 1.2 standard errors. On GDPval-AA, which Artificial Analysis runs, they scored 1844 and 1846 Elo at max effort: a tie.
The headline hides the effort scale. From the chart data on the launch page:
| Effort | Sonnet 5.5 | Cost per attempt | Opus 5.5 | Cost per attempt |
|---|---|---|---|---|
| Low | 20.0% | $0.76 | 38.5% | $1.29 |
| Medium | 28.8% | $0.83 | 57.6% | $2.94 |
| High | 43.0% | $1.94 | 64.2% | $3.88 |
| Xhigh | 61.5% | $5.30 | 66.4% | $7.35 |
| Max | 70.6% | $12.54 | 64.8% | $11.24 |
Sonnet 5.5's best point costs $12.54 an attempt against $7.35 for Opus 5.5's best, and Opus 5.5 at low beats Sonnet 5.5 at medium. Three caveats:
- Safeguards. Per the Sonnet 5.5 system card, a fallback model answered part of 10% of Opus 5.5's trials and 1.5% of Sonnet 5.5's. Anthropic says this "likely reduces" the Opus score.
- No outside confirmation. Neither model was on the public Terminal-Bench leaderboard on 2 October. Artificial Analysis's own run put them at 64% and 60%.
- GDPval-AA is 220 tasks from OpenAI's GDPval set, scored from blind pairwise comparisons. The live leaderboard now shows 1840 against 1846, with intervals of 24 and 23 points either way.
What changes for your stack now that Sonnet 5.5 is the claude.ai free tier model?
For a coding stack, not much, because Claude Code does not default to Sonnet 5.5 on any plan. Since version 2.1.280 on 22 September, it defaults to Opus 5.5 on every paid plan, Pro included. Sonnet 5.5 is what free claude.ai users get, per Anthropic's plan table and Simon Willison's test, although the launch page never says "free" or "default".
Anthropic's Sonnet page says "Anyone can chat with Claude using Sonnet 5.5 on Claude.ai", and the plan table gives the Free plan Sonnet and no Opus. Simon Willison ran a prompt against the free tier and wrote that Sonnet 5.5 is "now the model used for the free tier on claude.ai".
What matters more for a harness:
- Pro now starts on Opus. Before 2.1.280, Claude Code on Pro defaulted to Sonnet 5. The expensive model is now the default seat, and Sonnet is the one you assign, with
/model sonnet,opusplanor a subagent'smodel: sonnet. It needs Claude Code 2.1.284 or later. - Limits. Anthropic announced no usage limit change with Sonnet 5.5. Claude Code keeps separate Opus and Sonnet limits, so switching family keeps you working when one runs out.
- Haiku 5.5 is announced for "the coming weeks" with no date, price or model ID. Claude Haiku 4.5 is still the cheap seat.
On Stackness, as of 2 October 2026, 8 real profiles list Claude Code. Five of them list Claude Opus and one lists Claude Sonnet, on a profile that also has Opus (data sources). Tool pages are per family, so that is not a 5.5 count, and the numbers are small. Among the LLMs developers list on Stackness, it reads as one model per harness for now. If you run the split, list both.
Key numbers
- $2 / 10โ *โ *againstโ *โ *4 / 20โ *โ *permillioninputandoutputtokens,โSonnet5.5andOpus5.5,โwithcachereadsatโ *โ *0.20 on both (pricing, 2 October 2026).
- 35.40โ *โ *againstโ *โ *50.48: one Sonnet 5.5 Claude Code run and the same tokens at Opus 5.5 rates, a 30% saving, posted 29 September 2026 (r/ClaudeAI).
- 28.8% against 57.6% on Terminal-Bench 4.0 at medium effort, measured by Anthropic (launch page, 28 September 2026).
- 70.6% at max against 66.4% at xhigh on the same benchmark, at 12.54โ *โ *andโ *โ *7.35 per attempt.
- 1844 against 1846 Elo on GDPval-AA at max effort, run by Artificial Analysis.
- 8 real Stackness profiles list Claude Code, 5 list Claude Opus and 1 lists Claude Sonnet, as of 2 October 2026 (data sources).
Quick answers
Sonnet 5.5 vs Opus 5.5: which should I use for coding? Opus 5.5 for planning, review and open-ended work, Sonnet 5.5 for edits that are already scoped, at low or medium effort. At medium effort Opus 5.5 scores twice as high on Terminal-Bench 4.0, and Sonnet 5.5 costs less than a third as much per attempt.
Is Sonnet 5.5 half the price of Opus 5.5? Per token, yes, except cache reads, which cost $0.20 per million on both. In an agent session where cache reads are most of the bill, the saving on the same tokens is 20% to 30%.
Did Sonnet 5.5 beat Opus 5.5 on Terminal-Bench 4.0? At max effort, 70.6% against 66.4%, in Anthropic's own run. The lead is about 1.2 standard errors and costs $12.54 an attempt against $7.35.
What does opusplan do with the 5.5 models? It plans on Opus 5.5 and executes on Sonnet 5.5 on the Anthropic API. Each plan-mode toggle starts a fresh cache, and Sonnet 5.5 does not read Opus 5.5's thinking blocks.
Is Sonnet 5.5 the default in Claude Code? No. Claude Code defaults to Opus 5.5 on every paid plan since version 2.1.280. You select Sonnet 5.5 with /model sonnet, opusplan or a subagent setting.
Tools in this post
Use any of these tools?
Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.


