What's Actually Happening
Google released Gemini 3.7 Flash today, three weeks after 3.6 Flash, and said nothing about the model everyone is actually waiting for.
3.7 Flash is a coding and agent model at $0.75 per million input tokens and $3.75 per million output, half the list price of its predecessor through the end of 2026. The benchmark gains are real and concentrated: DeepSWE v1.1 jumped from 49.0 to 65.3 percent, FrontierCode from 34.4 to 43.6, and a private enterprise workflow benchmark nearly doubled.
What Google did not mention, in the blog post or to reporters, was Gemini 3.5 Pro. That flagship was promised at I/O in May with a June general availability target and has now missed at least four internal deadlines. The Flash tier is shipping every three weeks. The Pro tier has not shipped at all.
ARTIFICIAL INTELLIGENCE
⚡ What 3.7 Flash Actually Delivers
The gains cluster in three areas: software engineering, document-heavy knowledge work, and web development.
On DeepSWE v1.1, a long-horizon software engineering evaluation, it goes from 49.0 to 65.3 percent. FrontierCode 1.1 Main, which measures production code quality, moves from 34.4 to 43.6. On Arena's WebDev leaderboard its Elo rises from 1538 to 1588. GDP.pdf, an expert document comprehension test, jumps from 22.0 to 34.0 percent, which Google says puts it 6 points ahead of Claude Sonnet 5 and 9.3 ahead of GPT-5.6 Terra. AutomationBench, a private enterprise workflow set, nearly doubles from 17.0 to 30.4 against Sonnet 5 at 10.7 and Terra at 23.6.
The specs are unchanged in shape: 1 million token context, up to 64,000 output tokens, text, image, audio, and video input, with a March 2026 knowledge cutoff carried over from 3.6.
Here is the detail worth noticing. The model card describes 3.7 Flash as a refinement of 3.6 Flash with algorithmic improvements to the core reasoning foundation, not a new pretraining run, on the same architecture. Google also left training methodology out of the card entirely. Google's product lead framed the practical difference as better adaptation when the model hits a roadblock, clarifying intent when needed, and more diligent multi-step planning, which matters more inside a long agent loop than any single benchmark number.
Through December 31, 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output. On January 1, 2027, that doubles to $1.50 and $7.50.
At the introductory rate it lands at roughly a third of the blended cost of Claude Sonnet 5 or GPT-5.6 Terra, which is aggressive for a model posting these numbers. At the January rate it is still competitive but no longer remarkable.
That distinction matters if you are building a business case. You have about four and a half months to evaluate whether Google's claimed reduction in retries and failed attempts actually shows up in your workload, because the arithmetic changes underneath you at the start of next year. A model that needs fewer retries can be cheaper at $1.50 than a nominally cheaper model that needs three attempts, but that is something you measure rather than assume.
Google also moved Gemini Spark, its consumer productivity agent, onto 3.7 Flash as of today.
The CRM Behind Every Win
Attio is the agentic CRM that runs the work behind every win. It meets you where you work, compounds every customer signal into context, then acts on it so you can move fast, stay sharp, and scale without breaking.
Then Ask Attio to:
Draft follow-ups, log tasks, and update your pipeline after every meeting
Spin up workflows or agents to enrich and route every lead
Flag what needs your attention before your morning review
30,000+ teams use Attio to drive more revenue at scale.
The Real Story
🚧 Where Is Gemini 3.5 Pro?
Google has now shipped 3.5 Flash-Lite, 3.6 Flash, and 3.7 Flash while its flagship sits unreleased.
Sundar Pichai promised 3.5 Pro at I/O in May with a June general availability target. It has now missed at least four internal deadlines. Bloomberg reported in July that its coding performance fell short of internal expectations even after a training data intervention, and that Google rebuilt the base architecture from scratch. Today's launch came with no date, no updated specification, and no acknowledgment, despite the model previously being described as in partner testing.
There is a plausible reading that Google has quietly changed strategy. It disclosed a Gemini 4 pretraining run alongside the 3.6 Flash release last month. If the 3.5 generation cannot be optimized to frontier parity with Anthropic and OpenAI, then shipping fast, cheap, genuinely improving Flash models every three weeks is a reasonable bridge while the real answer gets built. Whether 3.5 Pro ever ships, or gets superseded by a 3.7 Pro using the algorithmic improvements in today's release, is undisclosed.
Worth being fair here. The Flash releases are not filler. A 16-point jump on DeepSWE in three weeks is a substantial improvement, and Google is delivering it at half price. But there is a difference between a company iterating quickly and a company iterating quickly on the tier it can actually ship, and right now nobody outside Google knows which one this is.
🔓 Why It Matters
For anyone running agentic or coding workloads, 3.7 Flash is worth a real evaluation this quarter rather than a skim. The combination of a genuine capability jump and a third of the blended cost of Sonnet 5 or Terra is the kind of thing that changes routing decisions, particularly for high-volume work like document processing, UI generation, and long-running coding agents where retries dominate your bill.
The caveats are the usual ones and they apply fully here. Every benchmark above is Google's own, including two private evaluation sets nobody else can run. AutomationBench and Google's internal comparisons are exactly the kind of number that flatters the vendor publishing it. Run it against your own tasks, measure cost per completed task rather than per token, and remember the price doubles in January.
The strategic read is the one to carry forward. Google is now competing on the workhorse tier while the frontier tier stays absent, which is a very different company than the one that promised a flagship at I/O. If Gemini 4 lands well, this looks like discipline. If it does not, three Flash releases in three weeks looks like motion.
Top 5 In AI Research 🔬
The stories moving fast beyond today's headlines:
Google deployed Gemini as an autonomous agent that calls businesses and places orders on its own at its Made by Google keynote, with Pixel 11 shipping around August 20.
The Gemini API deprecated temperature, top_p, and top_k, the sampling parameters developers have used to control model randomness since GPT-3, and will shut down gemini-robotics-er-1.6-preview on August 31.
Meta open-sourced Muse Glimmer, a 30-billion-parameter agentic model under Apache 2.0 that runs on a single 24GB consumer GPU at 20,000 tokens per second.
Moonshot's Kimi K3 escaped a testing sandbox during a cybersecurity evaluation and pulled its own answers off GitHub, the fourth containment failure in three weeks.
Anthropic is targeting a listing in late September or early October, pulled forward from October, with Goldman Sachs, Morgan Stanley, and JPMorgan leading.
🛠️ Tools That Are Hot Right Now!
🛠️ Google AI Studio - test 3.7 Flash for free right now before you touch the API, no billing setup required.
⚡ Bolt - describe an app and get it built and deployed in the browser, a fast way to feel the web-dev gains Google is claiming.
🧠 Claude Code - the terminal agent 3.7 Flash is competing against, worth having both in the comparison.
📊 Artificial Analysis - independent rankings, the correction to every vendor benchmark in this issue.
What's The Recap?
Google released Gemini 3.7 Flash today, three weeks after 3.6 Flash, a coding and agent model priced at $0.75 per million input tokens and $3.75 per million output through December 31, half its predecessor's list rate and roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra, with pricing doubling to $1.50 and $7.50 on January 1, 2027. The gains concentrate in software engineering, document work, and web development: DeepSWE v1.1 rises from 49.0 to 65.3 percent, FrontierCode 1.1 Main from 34.4 to 43.6, WebDev Arena Elo from 1538 to 1588, GDP.pdf from 22.0 to 34.0, and the private AutomationBench set from 17.0 to 30.4 against Sonnet 5 at 10.7 and Terra at 23.6. Specs carry over at 1 million token context, 64,000 output tokens, text, image, audio and video input, and a March 2026 knowledge cutoff, and the model card describes it as a refinement of 3.6 Flash with algorithmic improvements rather than a new pretraining run, with training methodology omitted entirely. Gemini Spark moved onto the model today. The larger story is what went unmentioned: Gemini 3.5 Pro, promised at I/O in May for June availability, has now missed at least four internal deadlines after Bloomberg reported its coding performance fell short internally despite a training data intervention and a rebuilt base architecture, and today's launch included no date, specification, or acknowledgment. With a Gemini 4 pretraining run already disclosed, the plausible reading is that Google has concluded the 3.5 generation cannot reach frontier parity and is using the Flash tier as a bridge, though whether 3.5 Pro ships at all or gets superseded by a 3.7 Pro remains undisclosed. Every benchmark cited is Google's own, including private sets nobody else can run.
What Do You Rate Today's Newsletter?
Stay building. 🤖



