In partnership with

What's Actually Happening

Google released Gemini 3.7 Flash today, three weeks after 3.6 Flash, and said nothing about the model everyone is actually waiting for.

3.7 Flash is a coding and agent model at $0.75 per million input tokens and $3.75 per million output, half the list price of its predecessor through the end of 2026. The benchmark gains are real and concentrated: DeepSWE v1.1 jumped from 49.0 to 65.3 percent, FrontierCode from 34.4 to 43.6, and a private enterprise workflow benchmark nearly doubled.

What Google did not mention, in the blog post or to reporters, was Gemini 3.5 Pro. That flagship was promised at I/O in May with a June general availability target and has now missed at least four internal deadlines. The Flash tier is shipping every three weeks. The Pro tier has not shipped at all.

The gains cluster in three areas: software engineering, document-heavy knowledge work, and web development.

On DeepSWE v1.1, a long-horizon software engineering evaluation, it goes from 49.0 to 65.3 percent. FrontierCode 1.1 Main, which measures production code quality, moves from 34.4 to 43.6. On Arena's WebDev leaderboard its Elo rises from 1538 to 1588. GDP.pdf, an expert document comprehension test, jumps from 22.0 to 34.0 percent, which Google says puts it 6 points ahead of Claude Sonnet 5 and 9.3 ahead of GPT-5.6 Terra. AutomationBench, a private enterprise workflow set, nearly doubles from 17.0 to 30.4 against Sonnet 5 at 10.7 and Terra at 23.6.

Gemini 3.7 Flash - X

The specs are unchanged in shape: 1 million token context, up to 64,000 output tokens, text, image, audio, and video input, with a March 2026 knowledge cutoff carried over from 3.6.

Here is the detail worth noticing. The model card describes 3.7 Flash as a refinement of 3.6 Flash with algorithmic improvements to the core reasoning foundation, not a new pretraining run, on the same architecture. Google also left training methodology out of the card entirely. Google's product lead framed the practical difference as better adaptation when the model hits a roadblock, clarifying intent when needed, and more diligent multi-step planning, which matters more inside a long agent loop than any single benchmark number.

Through December 31, 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output. On January 1, 2027, that doubles to $1.50 and $7.50.

Gemini 3.7/3.6 Flash Pricing

At the introductory rate it lands at roughly a third of the blended cost of Claude Sonnet 5 or GPT-5.6 Terra, which is aggressive for a model posting these numbers. At the January rate it is still competitive but no longer remarkable.

That distinction matters if you are building a business case. You have about four and a half months to evaluate whether Google's claimed reduction in retries and failed attempts actually shows up in your workload, because the arithmetic changes underneath you at the start of next year. A model that needs fewer retries can be cheaper at $1.50 than a nominally cheaper model that needs three attempts, but that is something you measure rather than assume.

Google also moved Gemini Spark, its consumer productivity agent, onto 3.7 Flash as of today.

The CRM Behind Every Win

Attio is the agentic CRM that runs the work behind every win. It meets you where you work, compounds every customer signal into context, then acts on it so you can move fast, stay sharp, and scale without breaking.

Then Ask Attio to:

  • Draft follow-ups, log tasks, and update your pipeline after every meeting

  • Spin up workflows or agents to enrich and route every lead

  • Flag what needs your attention before your morning review

30,000+ teams use Attio to drive more revenue at scale.

The Real Story

🚧 Where Is Gemini 3.5 Pro?

Google has now shipped 3.5 Flash-Lite, 3.6 Flash, and 3.7 Flash while its flagship sits unreleased.

Sundar Pichai promised 3.5 Pro at I/O in May with a June general availability target. It has now missed at least four internal deadlines. Bloomberg reported in July that its coding performance fell short of internal expectations even after a training data intervention, and that Google rebuilt the base architecture from scratch. Today's launch came with no date, no updated specification, and no acknowledgment, despite the model previously being described as in partner testing.

There is a plausible reading that Google has quietly changed strategy. It disclosed a Gemini 4 pretraining run alongside the 3.6 Flash release last month. If the 3.5 generation cannot be optimized to frontier parity with Anthropic and OpenAI, then shipping fast, cheap, genuinely improving Flash models every three weeks is a reasonable bridge while the real answer gets built. Whether 3.5 Pro ever ships, or gets superseded by a 3.7 Pro using the algorithmic improvements in today's release, is undisclosed.

Worth being fair here. The Flash releases are not filler. A 16-point jump on DeepSWE in three weeks is a substantial improvement, and Google is delivering it at half price. But there is a difference between a company iterating quickly and a company iterating quickly on the tier it can actually ship, and right now nobody outside Google knows which one this is.

🔓 Why It Matters

For anyone running agentic or coding workloads, 3.7 Flash is worth a real evaluation this quarter rather than a skim. The combination of a genuine capability jump and a third of the blended cost of Sonnet 5 or Terra is the kind of thing that changes routing decisions, particularly for high-volume work like document processing, UI generation, and long-running coding agents where retries dominate your bill.

The caveats are the usual ones and they apply fully here. Every benchmark above is Google's own, including two private evaluation sets nobody else can run. AutomationBench and Google's internal comparisons are exactly the kind of number that flatters the vendor publishing it. Run it against your own tasks, measure cost per completed task rather than per token, and remember the price doubles in January.

The strategic read is the one to carry forward. Google is now competing on the workhorse tier while the frontier tier stays absent, which is a very different company than the one that promised a flagship at I/O. If Gemini 4 lands well, this looks like discipline. If it does not, three Flash releases in three weeks looks like motion.

Top 5 In AI Research 🔬

The stories moving fast beyond today's headlines:

🛠️ Tools That Are Hot Right Now!

  • 🛠️ Google AI Studio - test 3.7 Flash for free right now before you touch the API, no billing setup required.

  • Bolt - describe an app and get it built and deployed in the browser, a fast way to feel the web-dev gains Google is claiming.

  • 🧠 Claude Code - the terminal agent 3.7 Flash is competing against, worth having both in the comparison.

  • 📊 Artificial Analysis - independent rankings, the correction to every vendor benchmark in this issue.

What's The Recap?

Google released Gemini 3.7 Flash today, three weeks after 3.6 Flash, a coding and agent model priced at $0.75 per million input tokens and $3.75 per million output through December 31, half its predecessor's list rate and roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra, with pricing doubling to $1.50 and $7.50 on January 1, 2027. The gains concentrate in software engineering, document work, and web development: DeepSWE v1.1 rises from 49.0 to 65.3 percent, FrontierCode 1.1 Main from 34.4 to 43.6, WebDev Arena Elo from 1538 to 1588, GDP.pdf from 22.0 to 34.0, and the private AutomationBench set from 17.0 to 30.4 against Sonnet 5 at 10.7 and Terra at 23.6. Specs carry over at 1 million token context, 64,000 output tokens, text, image, audio and video input, and a March 2026 knowledge cutoff, and the model card describes it as a refinement of 3.6 Flash with algorithmic improvements rather than a new pretraining run, with training methodology omitted entirely. Gemini Spark moved onto the model today. The larger story is what went unmentioned: Gemini 3.5 Pro, promised at I/O in May for June availability, has now missed at least four internal deadlines after Bloomberg reported its coding performance fell short internally despite a training data intervention and a rebuilt base architecture, and today's launch included no date, specification, or acknowledgment. With a Gemini 4 pretraining run already disclosed, the plausible reading is that Google has concluded the 3.5 generation cannot reach frontier parity and is using the Flash tier as a bridge, though whether 3.5 Pro ships at all or gets superseded by a 3.7 Pro remains undisclosed. Every benchmark cited is Google's own, including private sets nobody else can run.

Login or Subscribe to participate

Stay building. 🤖

Recommended for you

View all
caret-right