In partnership with

What's Actually Happening

Monday gave us three launches from three companies with three completely different ideas about what winning looks like.

Meta has a consumer app climbing the App Store faster than ChatGPT did. xAI has a new flagship model. Xiaomi has an open-weight model you can download for free.

One of them had a very good day. One of them had a fine day. And one of them woke up on Tuesday to find a Chinese phone company had matched its brand new model's benchmark score at roughly a fifth of the price.

ARTIFICIAL INTELLIGENCE
🚀 Muse Is Outrunning ChatGPT's Own Launch

Meta launched Muse on September 8. It is a personal AI agent led by chief AI officer Alexandr Wang, and it is structurally different from a chatbot. Each user's agent runs on a dedicated virtual machine in Meta's cloud with its own browser, and a separate system called Sentinel gates what it can reach, requiring approval before sensitive actions.

Meta Muse

Free tier, plus $20 and $100 monthly plans. No advertising.

The first twelve days, per third-party estimates:

1.8 million iOS downloads across the US and Canada, against ChatGPT's 1.3 million over the same stretch of its own launch.

2.8 million installs globally.

642,000 daily active users in the US, versus roughly 231,000 for ChatGPT at the comparable stage.

Number one on the US App Store, passing ChatGPT.

Now the number that explains all the others. Over 95 percent of Muse users already have Facebook, and 63 percent are on Instagram. Meta did not win a product fight this month. It pointed three billion existing accounts at a new icon.

xAI released Grok 4.7 on Monday, calling it its most capable model yet for coding and knowledge work. Bigger base model, extended reinforcement learning, better self-verification. Available through the Grok API, Cursor and Grok Build at $2 per million input tokens and $6 per million output.

Then the benchmarks came in.

Grok 4.7 Benchmarks Via SpaceXAI

On Artificial Analysis's Intelligence Index, Grok 4.7 scored 46. Claude Fable 5.1 and GPT-6 both scored 53.

On Terminal-Bench 4.0, which measures agentic coding, Grok 4.7 hit 26 percent. GPT-6 Astra hit 60. Claude Fable 5.1 hit 55. DeepSeek's V4.1 Flash, a free model, edged it out at 27.

The pricing is aggressive and the model is a real improvement on 4.6. But the gap to the frontier did not close this week, and on agentic coding specifically it widened.

200+ Claude Prompts Top Professionals Actually Use at Work

Claude can be your analyst, editor, and strategist.
But most professionals are using it to fix grammar.

These 200+ Claude prompts take it from grammar tool to your most powerful AI work assistant.

Sign up for Superhuman AI and get:

  • 200+ ready-to-use Claude prompts to get real work done in minutes — researched, tested, and used by professionals at Google, Microsoft, and NASA

  • Superhuman AI newsletter (4 min daily) so you keep learning new AI tools and skills to stay ahead in your career — the prompts are just the beginning

Within roughly a day of Grok 4.7, Xiaomi open-sourced the MiMo V2.6 series. Three models: Pro, Flash, and a distilled 9B.

MiMo V2.6 Pro is a mixture-of-experts model, 1.02 trillion total parameters with 42 billion active, 1M context.

On the same Artificial Analysis Intelligence Index, it also scored 46. Sixth overall, first among all open-weight models, a 20 point jump over V2.5.

Xiami’s MiMo

Same score as Grok 4.7. Here is the price:

Grok 4.7: $2.00 input, $6.00 output per million tokens.

MiMo V2.6 Pro: $0.435 input, $0.87 output.

That is roughly five times cheaper on input and seven times cheaper on output, for an identical index score. Xiaomi claims its costs run one twentieth to one sixtieth of leading international models at comparable intelligence.

And you do not have to take the API at all. Xiaomi published the weights, the technical report, the RL environments and the training code. You can run it yourself.

The individual benchmarks are mixed rather than a clean sweep. It posts 71.9 on DeepSWE v1.1 against Claude Opus 5's 74.0, and beats all three closed rivals on AutomationBench at 53.1. It gets buried on security reasoning, scoring 47.9 on ExploitBench where GPT-5.6 Sol scores 78.5.

⚔️ Three Bets, One Day

Read together, Monday was a snapshot of the three strategies left in this industry.

Meta is betting on distribution. It does not need the best model. It needs the model that is already installed on your phone, attached to the account you have had since 2009.

xAI is betting on price and speed of iteration. Ship often, undercut on cost, close the gap later.

Xiaomi is betting that the gap does not matter. If you can match a Western flagship's index score and hand out the weights, the frontier stops being a product and starts being a commodity.

The uncomfortable part for xAI is that the second and third strategies are the same strategy, and Xiaomi is running it better.

📚 What Can We Take Away?

For a year the open-weight conversation was about catching up. That framing is now out of date.

A Chinese consumer electronics company matched a US frontier lab's brand new flagship on a neutral index, one day later, at a fifth of the price, and gave the weights away. That is not catching up. That is arriving at the same place from a direction nobody was defending.

What this actually changes for you is the question you ask before you build. The old question was which model is best. The new question is what is the cheapest model that clears the bar for this specific job, because the answer is increasingly one you can self-host.

For Meta, the lesson runs the other way and is arguably more important. Muse is not winning on capability. It is winning because it launched into three billion accounts. Distribution has been the quiet advantage the whole time, and Meta just proved how much of a head start it buys.

Nobody won Monday outright. But the two companies that had the best day were the one that gave its model away and the one that never needed the best model to begin with.

Top 5 In AI Research 🔬

The stories moving fast beyond today's headlines:

  • ⚖️ Harvey, Abridge, Ramp and Rogo are moving off the frontier labs, adopting open-weight models or training their own to cut what they spend with OpenAI and Anthropic. Harvey moved its flagship product to Moonshot's Kimi K3. This is today's issue happening to real companies with real margins.

  • 💳 Shopify is putting Muse agents inside Shop Pay for one-tap agentic checkout across its storefronts, built on its Universal Commerce Protocol. It landed days after Amazon blocked Muse from its own store, which tells you the platform fight has already started.

  • 🕳️ Muse has a macOS zero-day, disclosed by Patrick Wardle, that lets local malware flip undocumented settings, redirect dictation traffic and lift authentication tokens. Worth knowing about before you install the app that just hit number one.

  • 🧮 OpenAI formed a math advisory group after its systems resolved more than 100 open problems in mathematics. Nine members, and a claim that will get scrutinized hard by people who do this for a living.

  • 🛑 Newsom signed an executive order on an AI kill switch on Friday, directing California to accelerate independent oversight and design emergency shutdown mechanisms for frontier models shown to resist being turned off.

🛠️ Tools That Are Hot Right Now!

  • 🐉 OpenRouter lets you point the same code at MiMo, Grok 4.7 and GPT-6 and compare cost per task directly. Given this week, that comparison is worth running.

    ⚙️ vLLM is how you actually serve an open-weight model like MiMo yourself instead of renting it.

    📊 Artificial Analysis is the neutral index both Grok 4.7 and MiMo were scored on. Worth bookmarking, since it is increasingly the scoreboard labs get judged by.

    🖥️ Cursor had Grok 4.7 available on day one, if you want to try it against your own codebase rather than a benchmark.

What's The Recap?

Meta's Muse is outrunning ChatGPT's launch numbers, which says more about owning three billion accounts than about the product. xAI shipped Grok 4.7 at an aggressive price and landed in the middle of the pack. Xiaomi open-sourced a model that matched Grok's index score for a fifth of the cost and published the weights.

The frontier is still Claude and GPT-6. Everything underneath it just got a lot cheaper.

Stay building. 🤖

Recommended for you

View all
caret-right