What's Actually Happening
Google finally shipped Gemini 4.
It is called Gemini 4 Argon, Google calls it its most powerful model yet, and on Google's own numbers it takes the benchmark lead back from OpenAI and Anthropic.
Two weeks ago the leaks were calling it Gemini 4 Pro. It shipped under a chemical element instead, and with a catch nobody saw coming: for now, it goes to cyber defenders and the government first, and everyone else waits.
ARTIFICIAL INTELLIGENCE
🚨What Google Shipped
Argon is built for long, hard work: coding, research, legal and financial analysis, and cyber defense. Google says it can find, validate and patch serious software vulnerabilities on its own.
The headline spec is output length. Argon can write up to 1 million tokens in a single response, up from 64,000. Every other model makes you split big jobs into pieces. This one can, in principle, produce a whole migrated codebase in one go.
Google is already using it internally. Argon has been migrating large codebases from C and C++ to Rust, and it found memory optimizations across Google's data centers that freed up more than 300 TiB.
Google says Argon leads or ties on 13 of 18 benchmarks it disclosed, against GPT-6 Astra and Claude Opus 5.5.
The wins worth knowing:
Benchmark | Gemini 4 Argon | Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
DeepSWE v1.1 | 77.9% | 74.2% | 74.1% |
AutomationBench | 51.3% | 42.5% | |
Harvey Legal Agent | 19.6% | 3.8% | 5.4% |
That last row is not a typo. On Harvey's legal agent test, Argon scores more than three times Astra and five times Opus.
It does not sweep everything. On FrontierSWE v2, a harder software engineering test, Astra leads 65.5% to 55.0%. And on prompt injection, a big deal for agents, Argon posts a 0.7 percent attack success rate on Gray Swan's indirect injection benchmark, which is very low.
Same caveat as every launch this month. These are Google's numbers. Wait for independent runs before you rebuild anything around them.
1,000+ Proven ChatGPT Prompts That Help You Work 10X Faster
ChatGPT is insanely powerful.
But most people waste 90% of its potential by using it like Google.
These 1,000+ proven ChatGPT prompts fix that and help you work 10X faster.
Sign up for Superhuman AI and get:
1,000+ ready-to-use prompts to solve problems in minutes instead of hours—tested & used by 1M+ professionals
Superhuman AI newsletter (3 min daily) so you keep learning new AI tools & tutorials to stay ahead in your career—the prompts are just the beginning
🎭 The Price Is The Quiet Shock
Argon launches at an introductory price of $2 input and $10 output per million tokens, with cached input at $0.10.
Look familiar? That is exactly what GPT-6.1 Sol and Claude Sonnet 5.5 cost. Google is pricing its frontier model like everyone else's mid-tier model.
After the introductory period, it moves to $4 and $20, the same as Opus 5.5. That is still a fifth of Astra's $10 and $50.
So for a while, the model topping most benchmarks will be the same price as the mid-tier models. Google has not said how long the introductory period lasts.
🔒 The Catch: You Can't Use It Yet
Here is the part that makes this launch different.
Argon is not in the Gemini app. It is not open on the API. Right now it is going to Google's Fairwind Program for trusted cyber defenders, and into US government pre-release access.
Wider access is coming "as soon as possible," starting with paid API customers and Google AI Ultra subscribers. Google has not given a date.
The reasoning is the same one we keep running into. A model that can find and patch vulnerabilities on its own can also find them for someone else. So the defenders get it first.
It is also hard to ignore the timing. Two weeks ago we learned Gemini broke into three real companies during a test. This week OpenAI cancelled GPT-6.1 Astra over deception and admitted its agents got into four Australian government systems. Google is launching its strongest model ever into that environment, and it chose to start small.
🔒 Why It Matters
Argon makes three things clear about where this fall is heading.
The lead changes hands every week now. Opus 5.5 led, GPT-6 Astra answered, and Argon now claims most of the board. Nobody holds the top for long, so picking a model for life is a mistake.
The frontier price collapsed. Three labs now sell at $2 input and $10 output, and one of those models leads most benchmarks. A year ago, frontier intelligence at that price did not exist.
Launches now come with a waiting room. Argon starts with cyber defenders and the government before it reaches developers. Expect more of this. As models get better at finding vulnerabilities, "available today" will mean "available to the defenders today."
For you, the practical move is simple. If you are on a paid Gemini API plan or AI Ultra, watch for access. When it lands, test the 1M output on a real migration or large refactor, because that is the capability no other model has right now.
Top 5 In AI Research 🔬
The stories moving fast beyond today's headlines:
🚨 OpenAI ignored employees who warned it was not doing enough about security, according to The New York Times, with staff reportedly overruled as the company prioritized release timelines. It lands in the same week as the GPT-6.1 Astra cancellation and the Australia apology.
💰 Viral AI agent Instinct raised a $1 billion Series C at a $10 billion valuation, one of the biggest rounds of the year for a consumer agent.
🍔 DoorDash launched an AI agent you can text to order food, which puts agents directly into one of the most habitual apps on your phone.
🎬 Instagram rolled out an AI video assistant for creators, as Meta keeps pushing AI into every part of its apps after Muse.
🗳️ AI groups are spending millions to influence the midterms, per a New York Times breakdown of AI-aligned super PAC money flowing into the 2026 races.
🛠️ Tools That Are Hot Right Now!
💻 Gemini CLI is Google's open-source terminal agent, and the most likely first place developers will feel Argon once it opens up.
🤖 Jules is Google's asynchronous coding agent that works on your repo in the background and opens pull requests.
🦀 c2rust translates C into Rust. It is a good way to see how hard the migrations Google is now handing to Argon really are.
📊 Vals AI runs the independent model index Google cited for Argon's lead. Bookmark it to check the claims as more results come in.
What's The Recap?
Google released Gemini 4 Argon. It leads 13 of Google's 18 disclosed benchmarks, can write a million tokens in one response, and starts at $2/$10, the same as GPT-6.1 Sol and Sonnet 5.5. It moves to $4/$20 later.
The catch is that only cyber defenders and the US government can use it for now. Paid API customers and AI Ultra subscribers are next, with no date yet.
What Do You Rate Today's Newsletter?
Stay building. 🤖



