Anthropic ships Claude Sonnet 5.5: Terminal-Bench 4.0 jumps from 10.3% to 70.6%, with 30%+ faster output and up to 30% lower cost per task
Synopsis
Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, positioned as a faster and cheaper model for everyday tasks and coding: Terminal-Bench 4.0 rises from 10.3% to 70.6%, GDPval-AA from 1449 to 1844 (near Opus 5.5's 1846), output generation is 30%+ faster, cost per task falls by up to 30% for most work, and it is the first Sonnet model to launch with cyber safeguards and anti-distillation classifiers.
Interpretation
Sonnet 5.5 substantially surpasses Sonnet 5 on agentic coding: 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3% and Opus 5.5's 66.4%; it scores 55.5% on CursorBench 4.0 (Sonnet 5: 34.1%; Opus 5.5: 57.8%). Relative to the previous Sonnet generation this is a generational jump rather than an incremental step, and its best CursorBench score lands within about two points of Opus 5.5. Drawn from multiple agentic coding benchmarks in the official results table, with evaluation details deferred to the Sonnet 5.5 System Card; the text also notes FrontierCode scores lower at Max effort than at Xhigh because of out-of-scope edits and a timeout.
On knowledge work and multidisciplinary reasoning, Sonnet 5.5 approaches Opus 5.5: 1844 on GDPval-AA v2.1 (Sonnet 5: 1449; Opus 5.5: 1846), 1811 on AA-Briefcase v1.1, 64.5% on Humanity's Last Exam with tools, 80.1% on OSWorld 2.1, and 61.6% on Chartography without tools. GDPval-AA spans 44 occupations and nine major industries, and Sonnet 5.5 sits about 400 points above Sonnet 5 while clearly outperforming Sonnet 5 and GPT-6 Sol on long-horizon knowledge work. Scores come from Artificial Analysis and Surge AI evaluations; a footnote states GDPval-AA and AA-Briefcase ran on a pre-release deployment with a bug that could degrade structured-output requests, which the company expects to have small effect and to understate performance, and which has since been fixed.
Cost and speed: pricing matches Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens, but fewer tokens are needed per task, cutting cost per task by up to 30% in testing while generating output 30%+ faster. The cost advantage comes from token efficiency and generation speed at an unchanged price, and on several benchmarks Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task. Based on the official pricing table and internal testing statements, accompanied by score-versus-cost-per-task charts across effort levels; the figures are the company's own measurements.
Safety and safeguards: on an automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty; because its cyber capabilities are comparable to Opus 5's, it is the first Sonnet model to launch with comparable cyber safeguards and fallbacks, plus the first with safety classifiers preventing reasoning extraction and expanded preserved thinking. Cyber safeguards and anti-distillation mechanisms previously reserved for the most capable models are extended to the Sonnet tier, while biology safeguards stay the same as Sonnet 5's. From the company's automated behavioral audit and newer containment evaluations; the text explicitly states that no set of evaluations reliably catches every failure and that Sonnet 5.5 may have tendencies not yet found.
Perspective
The release targets developers and knowledge workers who need fast iteration on everyday, well-scoped tasks: fixing bugs, understanding a codebase, and producing documents, slides, and spreadsheets, with effort levels trading cost and speed against quality (Medium is the default in Claude Code and the apps, High on the Claude Platform). The stated positioning is a lower-cost complement to Opus 5.5 rather than a replacement: Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Cyber safeguards target a narrow set of high-risk requests and visibly fall back to Sonnet 5, and biology safeguards target harmful requests, leaving routine software development and most life sciences work unaffected; teams needing more access can apply to the Cyber Verification Program and the Life Sciences Verification Program.
Readers should still watch that most scores are company-run, that some third-party results (Artificial Analysis, Surge AI) came from a pre-release deployment with a structured-output bug, and that GPT-6 Sol scores may not yet reflect the latest model version; FrontierCode scoring lower at Max effort than at Xhigh because of out-of-scope edits and a timeout shows effort and score are not monotonic. On alignment, the company states plainly that no set of evaluations reliably catches every failure and that Sonnet 5.5 may have tendencies not yet found; biology safeguards may flag some microbiology and virology requests in error. Migration also requires switching from thinking off to the between_tools setting, and moving conversations between accounts is affected by preserved thinking.
