Anthropic's new Sonnet arrives with an appealing promise: get ordinary work done faster without paying for the largest model. The release also arrives with enough integration changes to punish anyone who treats upgrading as a model-name swap. For builders, those two facts belong in the same conversation. A cheaper successful task is useful. A cheaper request that breaks the application, loses context or quietly changes which model answers is a different proposition.
Anthropic released Claude Sonnet 5.5 on September 28, positioning it for bounded coding tasks and document work. The company claims output generation is more than 30 percent faster than Sonnet 5 and that many tasks cost up to 30 percent less. Those are Anthropic's measurements, not independently reproduced results for this report. TechCrunch's launch coverage likewise attributes the speed and efficiency assertions to the company. The announcement establishes a new product and a testable performance claim, not a universal productivity dividend.
The published tariff explains the distinction. Anthropic's pricing documentation lists Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, the same rates as Sonnet 5. Cached reads cost $0.20 per million tokens. Five-minute cache writes cost $2.50 and one-hour writes $4 per million. A workload can therefore become cheaper through lower consumption or a different mix of billed activity without the token price falling. Calling this a blanket price cut would misdescribe what changed.
That also changes the comparison a buyer should make. I would measure cost per accepted task, including unsuccessful attempts, rather than compare the bill for a single answer. Define acceptance before testing: a code change passes the existing tests and stays within scope, or a document reconciles to its source material without missing required content. Track reviewer time separately. These are proposed purchasing criteria, not a claim about any customer's deployment. They keep an apparent saving from depending on work that a person quietly has to redo.
Anthropic reports a large Terminal-Bench 4.0 improvement, from Sonnet 5's 10.3 percent to Sonnet 5.5's 70.6 percent. It also says Opus 5.5 remains stronger for complex, open-ended work. One revealing launch footnote says Sonnet 5.5 scored worse at maximum than at extra-high effort on FrontierCode; in two examined cases, additional review led to a timeout or out-of-scope changes. The lesson is narrower than a model ranking: more activity can conflict with the acceptance criteria of the task.
Effort is the control that makes that tradeoff operational. Anthropic's documentation describes it as a behavioral setting affecting thinking, response text and tool calls, not a hard token allowance. Sonnet supports five levels, from low through maximum. A team should therefore test the setting and model together. Comparing a carefully tuned old configuration with a new default would not isolate the model change. It would compare two systems whose instructions about time, cost and thoroughness differ.
A useful effort sweep would include a repetitive extraction job, a bounded repair and a problem requiring several dependent actions. Keep the acceptance criteria fixed and record both successful and unsuccessful runs. If the lower setting holds quality for one class of work, route that class accordingly. If difficult tasks need more effort, account for it instead of averaging the expense out of sight. This is an evaluation design, not a promise that moving a slider produces a predictable percentage saving.
The migration guide identifies a change that can fail before quality is even tested. Sonnet 5.5 no longer accepts the disabled thinking setting; its lowest option is between_tools, which avoids up-front thinking but can produce progress updates between tool calls. That option works at low, medium and high effort, not the two higher levels. The guide also requires response handling by block type and preservation of thinking blocks in tool loops. Applications should not assume the first returned block is ordinary answer text.
Forced tool selection is another breaking change described in the guide. Requests specifying any tool or a named tool are rejected. Automatic selection is the replacement, with strict schemas where supported, but automatic selection can also produce a response without a tool call. That distinction matters for an application that previously relied on a call always occurring. Validation of a tool's arguments is not the same guarantee as execution of that tool. The application still needs to recognize an outcome in which no call happened.
Consider a support workflow that must look up an order before describing its shipping status. Under a proposed acceptance test, an answer without the required lookup should not count as a completed request simply because its language sounds confident. The team could stop, request the missing evidence or hand the task to a person according to its own policy. The important point is to preserve the business requirement while adapting to the new interface, not relax the requirement to make the migration dashboard look green.
Conversation handling creates a separate risk. The preserved-thinking documentation explains that returned thinking blocks are checked against compatible models and the conversation content that preceded them. Changes to earlier messages, tools or system instructions can invalidate blocks under the applicable enforcement rules. The documented responses include rejection or dropping invalid blocks. A request that proceeds after losing earlier reasoning is not identical to one that retained it, even though both may produce fluent output. This is why testing only a fresh, short conversation misses part of the migration.
That documentation also distinguishes a model switch from an edited conversation. Blocks the receiving model cannot read are removed from that request without being billed; the client is told to retain the full history rather than destructively rebuild it from what one model used. For a product with routing, the test should include a switch away and back, a saved-session restart and a long interaction that changes context. These are more revealing cases than asking two models the same isolated question and comparing their prose.
Safety outcomes need their own handling too. Anthropic's refusal documentation describes a successful HTTP response whose stop reason is refusal. It is not a transport failure, and partial streamed output should be treated as incomplete rather than delivered as a finished result. Some refusal categories can be billed before output, while mid-stream refusals can bill work already performed. That means the cost and completion record needs to distinguish a declined task from both a successful answer and a connection error.
The documented server-side fallback option is a beta on the Claude API, not a universally available feature across partner platforms. It can retry eligible declined requests, and the returned model identifies who actually answered. A deployment should record that identity along with the outcome. Otherwise, the team can believe it is measuring Sonnet 5.5 while evaluating a mixture of models. Supported fallback is also not a reason to weaken application-level permissions or keep retrying a prohibited task until something produces an answer.
Computer-use integrations deserve a similarly concrete review. Anthropic's current tool documentation describes a client toolset whose actions are executed by the customer's application in its own environment. A batch can contain multiple actions, which must run in order. If one fails, later actions should not execute, and every requested action still needs a corresponding result. A loop designed around a single action per response needs to be checked against that behavior rather than assumed compatible because a screenshot test worked.
The security guidance for that interface recommends isolation, limited privileges and human confirmation for consequential actions. It also warns that hostile instructions inside encountered content can still affect behavior. The practical implication is that improving a model does not justify expanding its permissions by default. If a workflow was authorized to inspect a document, a faster model should not suddenly be allowed to email it. Capability testing and authority design are separate review items, even when the same team owns both.
For a rollout, I would create a small migration gate with explicit failure categories: invalid requests, missing required actions, incomplete answers, lost conversation state and unexpected serving-model changes. Each category should have an owner and a reproducible test. Then evaluate quality and cost on requests that actually reached the intended configuration. This ordering prevents integration faults from being mistaken for evidence that the new model is unintelligent, and prevents polished answers from hiding a broken operational contract.
A limited production trial should preserve a known working configuration and a clear way to stop new traffic from reaching the replacement. Choose work with observable outcomes first, keep the review standard unchanged and expand only when those outcomes support it. For document work, that means checking the document, not the model's description of it. For code work, it means checking the resulting change in its intended environment. These are recommended release criteria, not reports that we have run a private benchmark or deployed Sonnet in a customer's system.
There is a real upside in getting this right. A model that meets the same acceptance bar with less waiting and less billed work can make a previously awkward workflow more practical. But the release-day record supports a bounded conclusion: Anthropic has shipped a new Sonnet, published efficiency claims and documented material interface differences. Independent launch coverage is not an independent reproduction of the full benchmark suite. Treat the release as an invitation to measure, with the migration details included in the experiment.
The best decision is not automatically to use the newest model everywhere or to keep the old one out of habit. It is to identify where the new configuration completes valuable work at a better total cost, then preserve the conditions that made that result true. Sonnet 5.5 may earn a substantial share of those jobs. Make it earn them on completed work, with intact context and appropriate authority, rather than on the speed of the first convincing answer.
LaunchPad positionMeasure cost per accepted task and test tool selection, conversation continuity and actual serving models before expanding a rollout.
This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.
