In partnership with

What's Actually Happening

Anthropic released Claude Opus 5.5 on Tuesday morning.

OpenAI released GPT-6 Sol and GPT-6 Luna minutes later.

Not the same week. Not the same afternoon. Minutes. Whatever release calendar these companies were running has collapsed into something closer to a reflex, and both of them cut prices on the way out the door.

Underneath the timing there is a real result, and it is not the one either press release is selling.

ARTIFICIAL INTELLIGENCE
🚨 What Anthropic Shipped

Opus 5.5, which Anthropic calls its new leading model and the strongest it has tested.

The price went down. $4 per million input tokens and $20 output, both 20 percent below Opus 5. Cache reads dropped 60 percent to $0.20. Anthropic puts the all-in saving at roughly 40 percent on typical workloads.

Pricing Via Anthropic

On Anthropic's own benchmarks, against Fable 5.1 and Opus 5:

Terminal-Bench 4.0: 66.4 percent, against 55.8 and 52.3.

CursorBench 4.0: 57.8 percent, against 51.8 and 46.6.

FrontierCode v1.1: 54.4 percent, against 50.3 and 48.0.

GDPval-AA v2.1: 1846 Elo, against 1735 and 1708.

Output is more than 30 percent faster. Anthropic also says it posted the best scores of any model to date on its automated behavioral audit and is much less likely to take hard-to-reverse actions. Sonnet 5.5 and Haiku 5.5 follow in the coming weeks.

Two models, split by job.

OpenAI’s Official Announcement

GPT-6 Sol handles complex work, coding and technical tasks. $2 input, $10 output per million tokens, down 50 percent from the 5.6 tier. OpenAI says it makes about half as many mistakes as its predecessor and reaches Astra-level reliability at much lower cost.

GPT-6 Luna is built for high volume clerical work, summarizing documents, pulling out information, answering straightforward questions. $0.10 input and $0.50 output. That is not a typo.

OpenAI credits the cuts to caching and inference improvements, and claims both models substantially outperform Anthropic's Fable and Opus models.

Sol is live in ChatGPT Work, Codex and the API. Luna is rolling out to the desktop app and to Free and Go users.

What does a $120k hire actually cost you?

A $120K salary might look straightforward on paper. But once you factor in employer taxes, statutory benefits, payroll, and country-specific employment costs, the number can look pretty different.

And those costs can vary significantly depending on where your employee is based. Before you build your hiring budget, or make an offer, it helps to know the full picture.

Oyster’s Global Hiring Cost Calculator lets you compare employment costs across 120 countries and estimate what a hire could really cost your business.

⚔️ Who Actually Wins

Here is where the press releases stop being useful, because OpenAI's comparison was against Opus 5, not the model Anthropic shipped that morning.

Run both new models on the same independent harness and the picture inverts.

Cost per task, measured on the Intelligence Index:

Effort

Opus 5.5

GPT-6 Sol

Low

$0.55

$0.13

Medium

$1.34

$0.25

High

$1.82

$0.37

Max

$5.98

$1.06

Quality, and this is the part OpenAI did not put in the announcement:

Sol's best possible result, 47.5 index points at max effort, still lands below Opus 5.5 running at its default medium setting, which scores 51.2.

Terminal-Bench 4.0: Opus 5.5 at medium hits 52.5 percent. Sol at max hits 43.9.

GDPval-AA: Opus 5.5 at medium scores 1576 Elo. Sol at max scores 1487.

AA-Briefcase: 1642 against 1483.

On business workflow automation they land in a dead heat, 61.2 against 61.7.

So Anthropic has the better model and OpenAI has the better deal, by a lot. Sol at medium costs about a fifth of what Opus 5.5 costs at medium. Whether that trade is worth it stops being a benchmark question and becomes a question about your specific workload.

One caveat worth holding onto. Anthropic's own Terminal-Bench figure is 66.4 percent while the independent run at medium effort reports 52.5. Both are real. They are measuring at different effort settings, and effort settings move these numbers more than most people realize.

The Part Nobody Is Saying Out Loud

On September 14, Dario Amodei published an essay called "We Must Pace the Frontier" calling for a global slowdown in AI development. He cited losing control of AI systems, cyberattacks and bioterrorism, and wrote that recursive self-improvement should be pursued very carefully, if at all.

Sam Altman, Demis Hassabis and Elon Musk all publicly agreed.

Eight days later, Anthropic shipped its most capable model ever, and OpenAI shipped two models within minutes of it.

There is a fair defense, and Anthropic makes it directly: Opus 5.5 posted the best behavioral audit scores of any model the company has tested, and is specifically less willing to take actions that cannot be undone. Pacing the frontier is not the same as stopping, and a safer strong model is arguably the point.

There is also the other reading, which is the one VC David Sacks offered when the essay landed. He argued it is simply good business for OpenAI and Anthropic to trade some raw power for reliability. Anthropic's annualized revenue has passed $100 billion, up from $65 billion in late July, with an IPO targeted for November at a valuation reported near $2 trillion.

Both readings can be true. But "we must pace the frontier" and "we shipped our strongest model in eight days" are going to sit in the same paragraph in a lot of articles this week, and the industry has not offered a version of that sentence that resolves cleanly.

🔓 Why It Matters

The interesting thing about today is not who won. It is that the question changed shape.

For two years the frontier was a ladder. One model was best, you paid for it, and the rest were compromises. Today two labs shipped within minutes, both cut prices, and the honest answer to which one you should use is now a question about your workload rather than a leaderboard position.

If you are doing agentic coding or hard knowledge work, Opus 5.5 is measurably ahead and you will feel it. If you are running high volume tasks where the bar is competence rather than brilliance, Sol at a fifth of the cost, or Luna at fifty cents a million output tokens, will beat paying premium rates for headroom you never touch.

The practical move is to stop picking a house model. Price both against your own evals, on your own tasks, with effort settings you actually plan to run in production. That number is the only benchmark that pays your bill.

And keep the speed itself in view. Two frontier labs released four models between them in under an hour, nine days after their CEOs agreed publicly that this should slow down. Whatever is governing the pace here, it is not the essays.

Top 5 In AI Research 🔬

The stories moving fast beyond today's headlines:

  • 🕳️ Plugin4Shell is a zero-click RCE hitting Claude Code, OpenAI Codex, GitHub Copilot and Gemini CLI through git hash collisions in the plugin supply chain. Millions of agent installs affected, and as of disclosure two of the four remained unpatched. Go check your versions.

  • 🔓 A three-person team used Claude Opus 5 to breach OpenAI, compromising employee accounts and reaching the internal monorepo. Opus 4.8 failed the same task, which makes this a capability jump, not a lucky break. OpenAI paid a $6,500 bounty.

  • 🦠 Cisco Talos disclosed CLOSEDQUORUM, described as the first fully autonomous multi-model AI command-and-control system that runs with no human operator, alongside CAIRN, a toolkit for detecting AI-integrated malware.

  • 🏛️ Trump said he will appoint an AI czar and create an "AI Force" modeled on Space Force, while dismissing AI safety warnings as a hoax and saying the new unit will cherish AI rather than stifle it. That is the policy backdrop the slowdown debate is now happening against.

  • 💰 Anthropic's annualized revenue has topped $100 billion, up from $65 billion in late July, with an IPO targeted for November at a valuation reported near $2 trillion. Useful context for every decision the company makes between now and then.

🛠️ Tools That Are Hot Right Now!

  • 🧪 promptfoo runs your actual prompts against Opus 5.5, Sol and Luna side by side and scores them. Given today, this is the single highest leverage hour you could spend this week.

  • 🔀 LiteLLM gives you one API across every provider, so swapping Opus 5.5 for Sol is a config change rather than a refactor. Price wars are only useful if you can move.

  • 📉 Langfuse traces cost and latency per request in production, which is how you find out what you are really paying per task instead of per million tokens.

  • 📊 OpenRouter is still the fastest way to put all four of today's models behind one key and compare them on a real workload before you commit.

What's The Recap?

Anthropic shipped Opus 5.5 and OpenAI answered with Sol and Luna minutes later, both cutting prices hard. Opus 5.5 is the stronger model, beating Sol's maximum effort while running at its default. Sol costs roughly a fifth as much, and Luna is priced in cents.

Eight days earlier, the CEOs of both companies agreed publicly that AI development should slow down. Draw your own conclusion about which of those two facts describes the industry better.

Login or Subscribe to participate

Stay building. 🤖

Recommended for you

View all
caret-right