Anthropic on Sept. 22 unveiled its new AI model, Claude Opus 5.5. Anthropic said it delivers performance similar to the higher-tier model Claude Fable 5.1 on many tasks while lowering typical workload costs by 40 percent versus its previous Claude Opus 5.
Claude Opus 5.5 is the first model in the new Claude 5.5 series. It is also the first model released since Chief Executive Dario Amodei (다리오 아모데이) on Sept. 12 published an essay calling for a slowdown in the pace of AI development.
The model is available from that day on Claude Pro, Max, Team and Enterprise plans. Anthropic raised the five-hour usage limits for Pro, Max, Team and seat-based Enterprise plans, and it provided subscribers with usage-limit resets that can be used at any time they choose.
Developers can use the model under the name "claude-opus-5-5". It also became available the same day through the Claude API as well as Amazon Bedrock, Google Cloud and Microsoft Foundry. GitHub also began rolling it out sequentially to paying GitHub Copilot users. Anthropic plans to release the small and mid-sized models Claude Sonnet 5.5 and Claude Haiku 5.5 within a few weeks.
API pricing is $4 per 1 million input tokens and $20 per 1 million output tokens. That is 20 percent lower than Opus 5. Cache read pricing, a major cost component for agent-type tasks and coding, was cut 60 percent to $0.2 per 1 million tokens. Accounting for a reduction in tokens used, typical workload costs under default settings are 40 percent lower than Opus 5. Output speed is more than 30 percent faster, and pricing for a "fast mode" that is up to 2.5 times faster is $8 per 1 million input tokens and $40 per 1 million output tokens.
Developer settings also changed. In Opus 5.5, thinking cannot be turned off, and reasoning depth is adjusted through an effort parameter. The default effort was lowered to medium from high in Opus 5.
On performance metrics, it led in agent-type coding, computer use and knowledge work. Its Terminal-Bench 4.0 score was 66.4 percent, higher than Fable 5.1 at 55.8 percent and OpenAI's GPT-6 Astra at 57.9 percent. It trailed GPT-6 Astra in some items in scientific research and work workflows. Anthropic said that, based on internal use, the gap with Fable 5.1 was not as large as the scores suggested.
Early use cases were also disclosed. One tester finished a code migration involving 680,000 lines in less than a day. In internal tests, it completed a task to port the load balancer HAProxy from C to Rust in 9.5 hours, shorter than with Fable 5.1, and at 51 percent lower cost. Anthropic said it also improved sentence readability by putting the most important information up front and reducing jargon and distinctive expressions.
In safety evaluations, Anthropic said both external testing results and its own judgment found that it did not exceed the threshold for the next capability level. U.S.-based METR concluded that AI research and development automation capability improved only slightly versus Fable 5.1 and that the likelihood of fully automating AI research and development was low. Anthropic also judged it did not exceed the next capability threshold under its responsible scaling policy.
Anthropic said stronger models that can fully automate AI research itself would require higher safety standards. It also said it does not believe current measures would meet that standard. Based on that judgment, it applied safeguards for Opus 5.5 at the same level as Fable 5.1.
In cybersecurity, it can find and fix bugs in code it wrote itself, but many other tasks are automatically switched to Claude Opus 4.8. In biology, it applied the same safeguards as Fable 5.1. Anthropic also signaled an expansion of its Cyber Verification program and the launch of a Biosciences Verification program for vetted organisations.
In an automated behavioral audit based on about 2,000 scenarios, it posted better overall results than recent Claude models. In a new evaluation measuring tendencies to cross containment boundaries, boundary-avoidance attempts were about 85 percent lower than Opus 5 and Mythos 5.1.
In evaluations without safeguards, there were cases of attempts to escape or tamper with a sandbox in 1.5 percent of instances, and a regression was confirmed in which it more easily followed malicious instructions pasted into prompts by users. Anthropic said its assessment of catastrophic-harm risk remained at "low," the same as in an August report.