What's Actually Happening
Six days after Opus 5.5, Anthropic shipped Claude Sonnet 5.5.
Normally the Sonnet release is the sensible follow-up: a bit slower, a bit dumber, a lot cheaper. This one outscored Opus 5.5 on one of the benchmarks developers care about most, and came within two points of it on knowledge work.
It costs half as much.
ARTIFICIAL INTELLIGENCE
🟠 What Anthropic Shipped
Price: $2 per million input tokens and $10 output, unchanged from Sonnet 5. Opus 5.5 is $4 and $20.
Speed: more than 30 percent faster output than Sonnet 5.
Cost per task: up to 30 percent lower, because it finishes jobs with fewer tokens and fewer tool calls. Anthropic says Sonnet 5.5 at low or medium effort can beat Sonnet 5's best score at roughly a tenth of the cost per task.
Early customers put numbers on it. Box saw it run 2.4 times faster with 12 percent fewer tokens. Lovable saw a third fewer tool calls and half as many shell executions.
It is live now on the API, AWS, Google Cloud and Microsoft Azure as claude-sonnet-5-5.
⚔️ Sonnet 5.5 vs Opus 5.5
Here is the scoreboard from Anthropic's own benchmarks:
Benchmark | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
Terminal-Bench 4.0 | 70.6% | 66.4% |
GDPval-AA v2.1 | 1844 | 1846 |
OSWorld | 80.1% | 81.8% |
CursorBench 4.0 | 55.5% | 57.8% |
FrontierCode 1.1 | 46.2% | 54.4% |
So is the headline true? Partly, and the part that is true matters.
Sonnet 5.5 beats Opus 5.5 on Terminal-Bench, which measures agents working in a real terminal: running commands, fixing builds, getting things done. It is effectively tied on knowledge work, two Elo points apart, and close on computer use.
Opus still wins clearly on the hardest pure coding tests. FrontierCode is an eight point gap, and that is where Opus earns its price.
The practical read: for agent loops, terminal work and everyday tasks, Sonnet 5.5 is now the default and Opus is the specialist you call for the hardest problems.
One caveat from last week still applies. These are Anthropic's numbers at Anthropic's settings, and independent runs at different effort levels can shift them a lot. Wait for third-party results before you migrate a production pipeline on one table.
The AI Work Handbook That Cuts Your Workday in Half
The 8-hour workday is becoming a 4-hour workday for people who know how to use AI.
Everyone else is still catching up.
This AI work playbook shows you exactly how to cut your work hours in half using AI.
Sign up for Superhuman AI and get:
50+ step-by-step AI tutorials to cut your workload in half — covering every part of your workday, from emails to strategy, used by 1M+ professionals at Google, Microsoft, and NASA
Superhuman AI newsletter (4 min daily) so you keep discovering new AI tools and skills to stay ahead in your career — the playbook is just the start
🎯 The Real Target Is GPT-6 Sol
Look at the price again. $2 input, $10 output.
That is exactly what OpenAI charges for GPT-6 Sol, launched last Tuesday. Same price to the cent.
Last week the argument was that Anthropic had the better model and OpenAI had the better deal. Sonnet 5.5 is Anthropic's answer to that: match Sol's price, and bring Opus-level results on agent work with it.
Sol still has Luna underneath it at ten cents per million input tokens, and nothing from Anthropic comes near that yet. But on the tier where most developers actually live, the price gap just closed.
🔒 Two Things Nobody Expected in a Sonnet
Frontier cyber limits. Anthropic says Sonnet 5.5's cyber capabilities are comparable to Opus 5, so it inherits restrictions that used to be reserved for flagships. Higher risk cyber requests visibly fall back to Sonnet 5 instead of running on the new model.
Classifiers that block reasoning extraction. This is the first Sonnet built to stop people pulling its reasoning out. The timing is not an accident. Researchers recently decoded 315,320 thinking blocks from 6,708 public agent traces across OpenAI, Anthropic and Google systems, and recovered 62 API keys, 33 passwords and seven private keys from what the models were thinking.
The second one is worth a pause if you log agent traces anywhere public. Your model's reasoning can contain your secrets.
🛑 Meanwhile, OpenAI Pulled Its Next Model
On the same day Anthropic shipped, The New York Times reported that OpenAI cancelled GPT-6.1 Astra, which was planned for October.
According to Saachi Jain, OpenAI's head of safety systems, the model regressed on alignment and showed higher levels of deception, not always telling the truth about the actions it did or did not take. It would also push beyond the scope of a task without user permission, including using external tools.
It was also more capable than the model it was replacing. OpenAI shelved it anyway.
That lands right before OpenAI's DevDay, which is tomorrow.
Top 5 In AI Research 🔬
The stories moving fast beyond today's headlines:
🧪 UK AISI found GPT-6 Astra attempts supply-chain attacks in 29% of simulations, against 6 percent for GPT-5.6 Sol and 0 percent for GPT-5.5. AISI notes the model's growing awareness that it is being simulated may complicate the result.
⏸️ OpenAI paused training its latest models after agents probed US government sites in unexpected ways this summer. OpenAI says it will resume only once it is confident it has additional safeguards in place.
🌍 AMD is buying Fei-Fei Li's World Labs for $8.2 billion, and Li joins AMD as executive vice president and chief scientist. A chipmaker buying a world-model lab says a lot about where AMD thinks the next wave of compute demand comes from.
💵 SoftBank sold $11.1 billion in junk bonds at record yields to fund its OpenAI commitment. It is the biggest high-yield corporate bond sale ever.
🌐 Bill Gates says a global AI deal will be harder than Cold War nuclear talks, arguing that AI does not come with the obvious shared threat that got the superpowers to the table.
🛠️ Tools That Are Hot Right Now!
💻 Claude Code is the fastest way to try Sonnet 5.5 on your own repo. Its Terminal-Bench lead is exactly the kind of work Claude Code does.
📊 ccusage reads your local Claude Code logs and shows what you are actually spending per day and per model. Run it before and after switching to see the real saving.
⚡ Zed is a fast Rust-based editor with Claude built in. It is a good fit for a model whose selling point is speed.
🧮 Braintrust runs evals on your own tasks so you can check whether Sonnet 5.5 really matches Opus for your work before moving production traffic.
What's The Recap?
Anthropic released Sonnet 5.5 at Sonnet 5's price. It is 30 percent faster, up to 30 percent cheaper per task, beats Opus 5.5 on Terminal-Bench and effectively ties it on knowledge work. It also costs exactly what GPT-6 Sol costs, and that is the point.
Opus still leads on the hardest coding. And on the same day, OpenAI shelved GPT-6.1 Astra for deceptive behavior, right before DevDay.Stay building. 🤖



