Claude Sonnet 5: The First Agentic Model That Can Code and Use Tools Like Opus

Claude Sonnet 5 was just released by Anthropic with an ambitious claim: this is the most agentic Sonnet model ever. It can make plans, use tools like browsers and terminals, and run autonomously at a level that only large, expensive models could handle a few months ago.

What’s interesting is that Sonnet 5’s performance is close to Opus 4.8 — Anthropic’s flagship model — but at a lower price. For developers who have been using Sonnet 3.5, 3.6, or 3.7 for coding and tool use, this is a significant upgrade.

What’s New in Sonnet 5?

From Anthropic’s announcement, there are several key improvements over its predecessor (Sonnet 4.6):

1. Much Better Agentic Performance

Sonnet 5 shows substantial improvement in four main areas:

  • Reasoning — ability to break down complex problems into executable steps
  • Tool Use — integration with browsers, terminals, and external APIs
  • Coding — sustained coding for multi-step software engineering work
  • Knowledge Work — handling messy technical contexts without losing track

What makes this different: Sonnet 5 can handle sustained coding and debugging in messy technical contexts. Previous models tended to lose track when tasks were multi-step and required lots of tool switching.

2. Cost-Performance Sweet Spot

From benchmarks Anthropic published, Sonnet 5 (orange line) is a strict improvement over Sonnet 4.6 (gray line). It covers a much wider cost-performance range compared to Opus 4.8 (yellow line).

At medium effort levels, Sonnet 5 gives substantially improved cost efficiency. At higher effort, its performance can match Opus 4.8 for some tasks. Between Sonnet 5 and Opus 4.8, users can adjust effort level to find the right balance between cost and performance.

3. Better Safety

Anthropic’s safety assessments found that Sonnet 5 has an overall lower rate of undesirable behaviors compared to Sonnet 4.6. It’s also generally safer to use in agentic contexts.

Evaluations show Sonnet 5 has a much lower ability to perform cybersecurity tasks compared to current Opus models. This is actually good news from a safety perspective — the model is powerful but not too dangerous in the wrong hands.

Availability and Pricing

Sonnet 5 is available across all plans:

  • Free and Pro plans — becomes the default model
  • Max, Team, and Enterprise — available as an option
  • Claude Code — directly integrated
  • Claude Platform — via API with model ID claude-sonnet-5

Pricing (introductory until August 31, 2026):

  • Input: $2 per million tokens
  • Output: $10 per million tokens

After August 31, 2026:

  • Input: $3 per million tokens
  • Output: $15 per million tokens

Still much cheaper than the Opus tier, but with performance that comes close.

What This Means for Developers

If you’ve been:

  • Using Sonnet 3.x/4.x for coding but often losing context on long tasks
  • Needing a model that can handle multi-step debugging without constant hand-holding
  • Wanting agentic capabilities but finding Opus too expensive for production use

Sonnet 5 might be the sweet spot you’re looking for. It provides a strong agentic execution layer for software engineering work — sustained coding, tool use, debugging — at a price that’s still reasonable for production deployment.

One note: Sonnet 5 is not a replacement for Opus in all use cases. For tasks that need the deepest reasoning or most sophisticated analysis, Opus is still king. But for 80% of everyday agentic coding and tool use cases, Sonnet 5 seems more than sufficient.

We’ll have to wait for independent benchmarks to confirm whether Anthropic’s claims hold up in real-world scenarios. But from early signals, Sonnet 5 looks promising.

For more context, check out Opus 4.8 Plans, Gemini 3.5 Executes and AI Wrote 80% in 10 Minutes.

Has anyone tried it? Share your experience in the comments.


Discover more from Susiloharjo

Subscribe to get the latest posts sent to your email.

Discover more from Susiloharjo

Subscribe now to keep reading and get access to the full archive.

Continue reading