In partnership with

What's Actually Happening

Three releases, three very different sizes.

OpenAI opened up the output of a model you can't use. Mistral built one of the biggest models anyone has promised to give away. And Google shipped one small enough to fit in your pocket.

Here's what each one is, and why it matters.

OpenAI published 722 mathematical manuscripts, grouped into 372 families, all produced by an unreleased internal frontier model.

Everything is on GitHub under Apache 2.0: the papers, the source files, and for many of them, Lean proofs. Lean is a language computers can use to check a proof line by line, so those results don't depend on anyone taking OpenAI's word for it.

The numbers behind it:

About 4,000 problems attempted.

Roughly three hours of ChatGPT Pro-level thinking per accepted result, on average.

Topics across number theory, complexity theory, geometry and mathematical physics.

Why release it now? Last month OpenAI said its systems had resolved more than 100 open problems, and mathematicians pushed back. An independent advisory group on mathematics and AI said it did not endorse testing advanced problems on closed models. Publishing everything is OpenAI's answer: here's the work, check it.

And it does need checking. Not every paper has a Lean proof, and OpenAI itself warns that "some of the unformalized results could have issues." Some manuscripts take on extremely famous problems. Until mathematicians review them, treat those as claims, not breakthroughs.

Mistral unveiled Mistral Large 4, and yes, it is officially nicknamed Le Chonk.

It's a 1 trillion parameter mixture-of-experts model with 49 billion active, multimodal in and text out, trained in about two months on 4,000 Nvidia Grace Blackwell GPUs across more than 160 languages.

The nickname is an inside joke. In June, the internet started joking that Mistral was secretly building a giant model. Co-founder Guillaume Lample says ML4 "could be viewed as an initial version of the idea," and hinted that bigger models could follow.

The pitch is efficiency. Mistral's VP of Science, Pierre Stock, says it trained on "two to three times less than our Chinese competitors."

On Mistral's own benchmarks it edges or ties the top Chinese open models:

Benchmark

Mistral Large 4

Best Chinese rival

DeepSWE v1.1

62%

GLM-5.3, 61%

Harvey Legal Agent

15%

Kimi K3, 12.9%

FinWorkBench

67%

DeepSeek V4 Pro, 67%

The catch: the weights aren't out yet. It's available through Mistral's API now, with open weights planned around October 27, after three weeks of safety testing, and under a custom Mistral license, not Apache. Mistral says it will work with "trusted partners and governments" so the weights are used to defend, not attack.

Mistral calls this a "third way" between closed American models and open Chinese ones. It's also the third serious Western open-model push in a week, after Reflection's Beam and Aleph Alpha's Kolibri.

Your AI prompts are leaving out 80% of what you're thinking.

When you type a prompt, you summarize. When you speak one, you explain. Wispr Flow captures your full reasoning (constraints, edge cases, examples, tone) and turns it into clean, structured text you paste into ChatGPT, Claude, or any AI tool. The difference shows up immediately. More context in, fewer follow-ups out.

Used by teams at Microsoft, Amazon, and Notion. Try Wispr Flow free on Mac, Windows, iPhone, and Android.

At the other end of the scale, Google released EmbeddingGemma 2, an open embedding model that runs entirely on your device.

Embedding models are the quiet workhorses of AI. They turn content into numbers so apps can search by meaning instead of keywords. They're what powers "find that photo of the receipt" or "search my notes for anything about the launch."

What's new in version 2:

It handles text, code, images, audio and video, all in one shared space. You can search your voice memos with a text query.

It's tiny. 740 million parameters. Text-only uses about 191MB of RAM, and all modalities about 567MB.

It works offline. No network needed, so sensitive data never leaves the phone.

It's Apache 2.0 and supports more than 100 languages, with an 8,192 token context. That's enough for 29 images, 58 video frames or about five and a half minutes of audio.

One note for builders: Google says it ships with no safety tuning and no output moderation, so that part is on you.

🔓 Why It Matters

Put the three side by side and you get the three directions AI is moving at once.

Up. Mistral is proving you don't need to be American or Chinese to build a trillion-parameter model, and you don't need the biggest cluster either.

Down. Google is putting multimodal search on a phone with no internet. A lot of everyday AI will run like this: small, local and private, with no API bill.

Out. OpenAI is showing what frontier models can do before you're allowed to use them, and it's doing it in a form anyone can check. Lean proofs make "trust us" unnecessary.

For builders, the most useful one today is the smallest. If your app has search, EmbeddingGemma 2 is a free, offline upgrade worth testing this week. Mistral's weights come in three weeks. The math papers will take mathematicians months to judge.

Top 5 In AI Research 🔬

The stories moving fast beyond today's headlines:

  1. 🛡️ Anthropic expanded its Cyber Verification Program into three tiers and folded in Project Glasswing. Anthropic says the work found more than 129,000 verified software flaws between April and July, over 33,000 of them rated critical or high.

  2. ⚡ Google signed a 3.6 gigawatt power deal with Constellation Energy, including a 20-year agreement for 890 MW of nuclear, to keep its data centers running.

  3. 🤝 Meta and Sierra published the Personal Agent Protocol, an open standard for AI agents in commerce, with Walmart, Shopify and Stripe among the founding partners.

  4. 🌙 Moonshot AI hit a $50 billion valuation and is reportedly targeting a Hong Kong IPO in early 2027. The company behind Kimi is going public.

  5. 🇰🇷 South Korea plans a $3.5 billion frontier AI program starting next year, aiming for parity with leading Chinese open models.

🛠️ Tools That Are Hot Right Now!

📐 openai/math on GitHub is the full release. If you know any Lean, the formalized proofs are the most interesting part.

🔎 Sentence Transformers is the easiest way to drop an embedding model like EmbeddingGemma 2 into your own search or RAG pipeline.

🐱 Mistral Le Chat is where to try Mistral's models in the browser while you wait for Le Chonk's weights.

What's The Recap?

OpenAI put 722 math papers from a model you can't use on GitHub, Mistral announced a trillion-parameter model whose weights arrive around October 27, and Google shipped a 740 million parameter embedder that does multimodal search offline on a phone. The takeaway is that model size and model usefulness have split apart: the biggest model isn't out, the most impressive one may never ship, and the smallest is the only one you can build with tonight. So use a simple filter for any launch: can you run it, can you check it, and what does it cost? Today that means testing EmbeddingGemma 2 if your app has search, putting October 27 in your calendar for Le Chonk, and treating the math papers as claims until a Lean proof or a mathematician says otherwise.

Login or Subscribe to participate

Stay building. 🤖

Recommended for you

View all
caret-right