Anthropic Releases Claude Sonnet 5.5 at 70.6% on Terminal-Bench
Anthropic launches closed-weights Claude Sonnet 5.5 with a 1M-token context window, 70.6% on Terminal-Bench 4.0, and flat $2/$10 pricing.

Stock photo for illustration only, not from the actual event
- Claude Sonnet 5.5 is now live on the Claude Platform, AWS, Google Cloud, and Microsoft Azure
- Supports a 1M-token context, 128K max output, and June 2026 reliable knowledge cutoff
- Achieves 70.6% on Terminal-Bench 4.0 with adaptive thinking enabled by default
- Maintains list pricing at $2 per million input tokens and $10 per million output tokens
Anthropic has officially rolled out its latest artificial intelligence model, Claude Sonnet 5.5. The new model is currently live and accessible on the Claude Platform, alongside major cloud services including AWS, Google Cloud, and Microsoft Azure. Because it is structured as a closed-weights model, self-hosting is not available for users.
According to the model overview specifications, Claude Sonnet 5.5 supports a massive 1M-token context window, a maximum output of 128K tokens, and features a reliable knowledge cutoff set to June 2026. Furthermore, adaptive thinking is turned on by default, with effort levels spanning across five distinct tiers: low, medium, high, xhigh, and max.
All benchmark scores provided are vendor-reported by Anthropic from the official launch post. The Max evaluation result came in lower than the Xhigh tier because at the Max level, the model frequently executed multi-agent code reviews. This behavior occasionally resulted in timeouts or out-of-scope edits, which the FrontierCode system penalizes. Additionally, Anthropic noted that Opus 5.5 remains clearly superior for complex and open-ended assignments.

Stock photo for illustration only, not from the actual event
The release of Claude Sonnet 5.5 highlights Anthropic's strategic shift toward improving token efficiency rather than reducing baseline sticker prices. Real-world cost reductions are achieved because the model requires fewer tokens and tool calls per task, allowing enterprise users to lower total operational expenditures based on verified customer metrics.
List pricing remains completely unchanged from Sonnet 5, holding steady at $2 per million input tokens and $10 per million output tokens. Cache reads are priced at $0.20 while cache writes cost $2.50 per million tokens. These rates stand at precisely half the cost of Opus 5.5, which charges $4 and $20, demonstrating that financial savings stem entirely from structural token efficiency rather than a direct price cut.
Customer deployment data strongly supports these efficiency gains. Balyasny Asset Management recorded approximately 121K tokens per answer compared to 497K on Sonnet 5. Meanwhile, Base44 reported 3.6 iterations per application build contrasted with 7.7 iterations for Opus 5, and Zendesk successfully processed support tickets 20% faster.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment