Skip to content
AI ConnectPowered by VELENTIS
AI-generated2 min

Anthropic Releases Claude Sonnet 5.5: Significant Leap in Agentic Coding Benchmarks

Anthropic launched Claude Sonnet 5.5, a faster workhorse model scoring 70.6 percent on Terminal-Bench 4.0, directly replacing the previous base model in claude.ai's free tier.

This article was AI-generated and published automatically. Context, labelling and all sources at the end of the article.

(KI-generiertes Symbolbild: Gemini / AI Connect)

Just six days after deploying its flagship Opus 5.5, AI developer Anthropic officially rolled out Claude Sonnet 5.5 on September 28, 2026. The release marks a deliberate division of labor across the company's product lineup. While Opus remains targeted at complex, high-tier computational challenges, Sonnet is designed as an efficient, highly scalable workhorse for day-to-day production environments in engineering teams and enterprises.

The most pronounced performance improvements appear in agentic software development tasks. On the Terminal-Bench 4.0 benchmark, Sonnet 5.5 jumped to a success rate of 70.6 percent, up from 10.3 percent achieved by Sonnet 5. This score also places Sonnet 5.5 ahead of Anthropic's flagship Opus 5.5, which recorded 66.4 percent on the same test. The results illustrate an industry pattern where compact, task-focused architectures can outpace larger models on practical coding duties.

Performance gains were also reported across broader knowledge tasks and multimodal control. In the knowledge-work evaluation GDPval-AA v2.1, Sonnet 5.5 registered an Elo rating of 1844. Furthermore, Anthropic highlighted computer-use capabilities, noting that Sonnet 5.5 is the first model in its lineage to autonomously complete the video game Pokémon Red solely using screenshot inputs and direct desktop interaction.

Beyond benchmark milestones, the update emphasizes operational efficiency. According to Anthropic, Sonnet 5.5 runs approximately 30 percent faster than Sonnet 5. At the same time, actual operational costs per task decrease by up to 30 percent, largely driven by refined token management. For engineering organizations, these architectural tweaks translate into lower execution latencies and diminished resource expenditures across repetitive pipelines.

Anthropic set commercial API pricing at 2.00 dollars per million input tokens and 10.00 dollars per million output tokens. Consumer access has also changed immediately, as Sonnet 5.5 replaces the previous default base model in the free tier of claude.ai. This rapid deployment provides broader developer access to state-of-the-art agentic workflows without requiring premium tier subscriptions.

What this means for you

Sonnet 5.5 offers software teams a meaningful reduction in operating expenses for agent-driven coding tasks. Because it surpasses Opus 5.5 on Terminal-Bench 4.0, engineering departments can safely substitute expensive flagship calls with the more economical Sonnet tier.

Perspectives

Coverage: 1× US · 2× Other

One story, several angles: how each source frames the topic, each with a verbatim quote.

Leaning: 1× Vendor PR

  • anthropic.comVendor PRUS

    Anthropic positions Claude Sonnet 5.5 as a faster, lower-cost partner to Opus 5.5, highlighting a dramatic performance leap in agentic coding evaluations.

    Original quote

    „Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%.“

    anthropic.com
  • vellum.aiOther

    Vellum breaks down the benchmark results in detail, emphasizing that Sonnet 5.5 outperforms even the flagship Opus 5.5 in command-line agentic coding.

    Original quote

    „Terminal-Bench 4.0 measures how effectively an autonomous agent executes complex, multi-step engineering tasks inside a live command-line environment.“

    vellum.ai
  • simonwillison.netOther

    Simon Willison reviews the release from a developer's perspective, emphasizing its strong coding capabilities alongside the strategic advantage of offering it in Claude's free tier.

    Original quote

    „Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks“

    simonwillison.net

Source classification is maintained editorially (political spectrum only where consensus is broad; vendor communication is PR, not journalism). Unlabelled sources are unclassified: we do not guess.

Evidence

Well sourced
78/100
  • Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Opus 5.5.

    verified
  • Sonnet 5.5 scored 70.6 percent on Terminal-Bench 4.0, up from 10.3 percent for Sonnet 5 and beating Opus 5.5 at 66.4 percent.

    verified
  • The model runs roughly 30 percent faster than Sonnet 5 and cuts per-task costs by up to 30 percent.

    verified
  • API pricing is set at 2.00 dollars per million input tokens and 10.00 dollars per million output tokens, replacing the free tier base model on claude.ai.

    single source

The evidence score is computed, not hand-set: from confidence, the number of sources and the share of verified statements.

Source & transparency

As of: September 30, 2026

AI-generatedAI-generated: produced automatically from vetted sources with technical quality checks (source, quote and figure verification); no human sign-off of each item before publication

Sources
3
Verified statements
3 / 4
Evidence score
78Well sourced

Want to put this into practice?

We connect you with suitable AI providers from the DACH region, free of charge and without obligation.

What's next?