Claude Sonnet 5
Sonnet just reached for Opus's shoulder.
Same price as before — $3 per million input tokens, $15 output. Yet on a benchmark for autonomous terminal-driving agents, it beat Opus 4.8 outright. The role Anthropic's mid-tier model plays is quietly being rewritten.
What Changed
Same price. A much bigger leap.
The third Sonnet generation just crowded into Opus's territory.
On June 30, 2026, Anthropic officially announced Claude Sonnet 5, the fifth generation of the Sonnet line. It's built for coding, autonomous agent work, and professional knowledge tasks, and its API pricing holds at the same $3 per million input tokens and $15 per million output tokens as its predecessor, Sonnet 4.6 (an introductory $2/$10 applies through August 31, 2026).
Until now, the Sonnet line sat beneath Opus 4.8 and Fable 5 as the cost-conscious workhorse model. But as TechCrunch reported, Sonnet 5 scored 80.4% on Terminal-Bench 2.1 — a benchmark for autonomous agents driving a terminal — beating Opus 4.8's 74.6%. It's an unusual reversal: a mid-tier model pulling ahead of its flagship sibling on a specific metric.
By the Numbers
Where it actually beats Opus
Three benchmarks that matter directly for agentic work.
Reports also say Sonnet 5 edges out Opus 4.8 on GPQA-AAA v2, a knowledge-work reasoning benchmark. On the other hand, on SWE-bench Pro — the hardest coding benchmark — it scores 63.2%, a big jump from Sonnet 4.6's 58.1% but still short of Opus 4.8's 69.2%. This isn't "Opus beaten across the board" — it's a reversal concentrated specifically in agentic operation and knowledge work.
Three Generations
Three generations, side by side
| Sonnet 4.6 (previous) | Sonnet 5 (new) |
|---|---|
| Terminal-Bench 2.1: 67.0% | 80.4% (beats Opus 4.8's 74.6%) |
| SWE-bench Pro: 58.1% | 63.2% (Opus 4.8: 69.2%) |
| OSWorld-Verified (computer use): 78.5% | 81.2% |
| API pricing: $3 / $15 per million tokens | Holds at $3 / $15 (intro $2/$10 through Aug 31) |
In its official announcement, Anthropic frames Sonnet 5 as "a substantial upgrade" on the agentic qualities developers care about most — reasoning, tool use, coding, and knowledge work.
Who Benefits
Who this actually helps, and how
Engineers — more trust in terminal-driving agents
An 80.4% on Terminal-Bench 2.1 means Opus-class or better reliability for autonomous agents operating a CLI or shell, at Sonnet pricing. Beyond Claude Code, the model is also selectable through Cursor, VS Code, and GitHub Copilot.
Business leaders — Opus-level judgment without a price hike
Without raising API spend, teams get accuracy on knowledge-work evaluations that slightly edges Opus 4.8. Moving bulk document processing and decision-support workloads from Opus to Sonnet 5 can raise quality without raising cost.
PMs and product owners — the default model just switched under you
On Free and Pro plans on claude.ai, Sonnet 5 is now the automatic default. If internal tooling pins a model ID, switching it to claude-sonnet-5 is enough to pick up the gains. Workloads already running cheaply on Haiku 4.5 see little benefit from this release.
Next Steps
What to do next
First, point any API calls or agent configs still pinned to Sonnet 4.6 at claude-sonnet-5 and evaluate against real workloads. Second, validate the cost advantage on bulk workloads while the introductory $2/$10 pricing still applies, through August 31. Third, don't force a switch on low-cost, routine work already handled well by Haiku 4.5 — the biggest gains here are concentrated in terminal- and computer-driving agents and knowledge-work-leaning tasks.
Caveat
Not an unqualified win
There are reasons for caution. Reports say Sonnet 5 uses a new tokenizer that produces roughly 30% more tokens than Sonnet 4.6 for the same text. Even with the per-token price unchanged, the actual bill can rise. And on SWE-bench Pro, the hardest coding benchmark, it still sits at 63.2% against Opus 4.8's 69.2% — a meaningful gap. Rather than handing every task to Sonnet 5 outright, the hardest problems still call for the flagship model.