Anthropic Releases Claude Sonnet 5.5 with 30% Speed Boost and Lower Task Costs

앤트로픽이 AI 모델 '클로드 소넷 5.5'를 공식 출시했다. 이전 모델 대비 출력 속도를 30% 이상 높이고 작업 처리 비용을 최대 30% 절감한 것이 특징이다. API 가격과 주요 성능 평가, 안전장치 세부 내용

작성자

카테고리:

Anthropic officially released its AI model ‘Claude Sonnet 5.5’ on September 28, 2026. The model increases output speed by over 30% compared to Sonnet 5 and cuts task processing costs by up to 30%. Anthropic aims to capture everyday workflows and coding demand by delivering faster execution and a lighter operational footprint than its top-tier flagship.

Key Takeaways

  • Anthropic officially launched its AI model ‘Claude Sonnet 5.5’ on September 28, 2026.
  • Output speed improved by more than 30% compared to its predecessor, with task costs reduced by up to 30%.
  • API pricing is set at $2 per million input tokens and $10 per million output tokens, lowering total costs by reducing token consumption.
  • It scored 70.6% on the Terminal-Bench 4.0 coding benchmark, surpassing the higher-tier Opus 5.5.
  • Detailed comparisons with ChatGPT for end users and third-party validation of enterprise test results remain unconfirmed.

Maintaining Token Unit Prices While Cutting Token Consumption

Maintaining API Unit Prices While Reducing Token Consumption

API pricing for Claude Sonnet 5.5 is set at $2 per million input tokens and $10 per million output tokens. This matches the pricing for Sonnet 5 and OpenAI’s ‘GPT-6 Sol’.

The two companies took different approaches to cost efficiency. OpenAI cut the token unit price of GPT-6 Sol by 50% compared to GPT-5.6 Sol. In contrast, Anthropic maintained token unit prices while lowering practical task expenses by reducing the total tokens and tool calls needed to finish a task.

ModelInput Price / 1M TokensOutput Price / 1M TokensCost Reduction Approach
Claude Sonnet 5.5$2$10Reduced token consumption and tool calls
GPT-6 Sol$2$1050% cut in token unit price

Related Articles
OpenAI and Anthropic Simultaneously Launch Low-Cost AI Models, Fueling Price Competition
U.S. Enterprises Pivot to Open-Weight Models Over Costly Proprietary AI, Reaching 56% Token Share

70.6% on Command-Line Coding: Key Benchmark Results

70.6% on Command-Line Coding: Key Benchmark Results

Sonnet 5.5 achieved 70.6% on Terminal-Bench 4.0, a benchmark assessing multi-step command-line coding capabilities. This significantly surpasses Sonnet 5’s 10.3% and edges out the premier Opus 5.5 at 66.4%.

On GDPval-AA v2.1, which measures real-world professional and industry tasks, Sonnet 5.5 scored 1,844 points, nearing Opus 5.5’s 1,846 points. It recorded 80.1% on OSWorld 2.1 for computer control capabilities (Opus 5.5: 81.8%) and 61.6% on Chartography for chart interpretation (Sonnet 5: 15.6%).

In the FrontierCode coding benchmark under high reasoning effort, Sonnet 5.5 scored approximately 10 points higher than Sonnet 5. Task costs incurred during the evaluation ran at approximately 10% and 7% of Sonnet 5’s levels.

New Cybersecurity Safeguards and Deployment Channels

Anthropic also enhanced safety features. For the first time in the Sonnet family, dedicated cybersecurity safeguards are built in to route high-risk cybersecurity requests to Sonnet 5 for handling.

An anti-distillation safety classifier was added to block attacks that harvest large response volumes through fake accounts to copy model behavior. Automated behavioral audits across roughly 1,850 scenarios showed alignment, misuse prevention, and honesty scores matched or exceeded Sonnet 5.

Sonnet 5.5 is available through the Claude app, Claude Code, Claude platform, and Anthropic API, as well as on Amazon Web Services (AWS), Google Cloud, and Microsoft Azure.

Unverified Claims and Current Limitations

Detailed real-world use case comparisons with ChatGPT paid tiers from a general consumer standpoint were not confirmed in the available materials.

Enterprise performance metrics—such as Balyasny’s 76% token reduction in financial workflows, Slack’s 14% drop in output tokens, and Zendesk’s 20% faster ticket handling—are self-reported by Anthropic and lack independent third-party verification.

Additionally, external cross-validation for certain benchmark figures, such as the 61.6% on Chartography and the 10-point gain on FrontierCode, along with speculation regarding launch timing ahead of an IPO, have not been officially verified.

Related Articles

출처

1-Minute Summary Video

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다