In partnership with

What's Actually Happening

For six days a free model called Ox Alpha sat near the top of third-party usage charts while nobody knew who made it.

On August 26, Zhipu ended the guessing. Ox Alpha is GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with only 18 billion active, the first natively multimodal release in the GLM-5 series, priced at $0.15 per million input tokens, with MIT-licensed weights that went to Hugging Face the same day.

The community had already guessed the exact name days earlier, working backwards from a clue: GLM-5.3 shipped text-only on August 14, but Ox Alpha was accepting images and video, so it had to be an unreleased multimodal sibling.

And here is the part that makes it more than a launch. Zhipu's flagship GLM-5.3 was announced two weeks ago as the strongest open-weights coding model available. Its weights still are not out. The cheaper sibling shipped its weights on announcement day.

ARTIFICIAL INTELLIGENCE
🐉 What It Actually Is

GLM-5.3-Flash runs 320 billion total parameters with 18 billion active across 45 layers. That active footprint is roughly half of GLM-5.2's, which is where the cost savings come from. Zhipu says it outperforms GLM-5.2 across its reported coding and agentic tests despite running a fraction of the active parameters, and approaches Claude Opus 4.8 on its internal coding benchmark.

It is the first GLM-5 model to natively take text, image, and video input, rather than routing vision through a separate adapter bolted on afterward. Context runs to one million tokens, with long-context serving cost that Zhipu puts at roughly a third of GLM-5.3's.

Zhipu also describes it as the first open-source frontier model using a hybrid sparse-attention plus linear-attention design, paired with something it calls Manifold-Constrained Hyper-Connections for scaling and an IndexPool mechanism that compresses four indexer key vectors into one weighted pool to cut long-context latency. Pretraining ran on 30 trillion multimodal tokens.

The pricing is the headline for most people. $0.15 per million input tokens is a tenth of the flagship GLM-5.3 rate, or a twentieth during a limited-time discount. GLM Coding Plan users also get triple their GLM-5.3 quota.

Zhipu says the whole thing runs on Chinese AI chips, which if accurate is another data point that export controls have a ceiling on what they can prevent.

🕵️ The Stealth Playbook

The way this was released is arguably more interesting than the model.

Zhipu shipped it anonymously as Ox Alpha on a routing platform starting August 20, free, with no announcement. Developers used it heavily for six days. Independent benchmarks accumulated. Community speculation worked out both what it was and what it would be called, correctly, before the company said a word.

This is not the first time. In February a codename called Pony Alpha appeared on OpenRouter, processed over 40 billion tokens on its first day, topped the platform's popularity rankings, and was then confirmed as GLM-5, Zhipu's 744-billion-parameter flagship.

The logic is sound and slightly uncomfortable for everyone else. Launch-day hype distorts early impressions in both directions, since people either arrive primed by a benchmark chart or dismiss it as vendor marketing. Ship anonymously and what you get instead is real usage data from developers who had no idea whose model they were testing, and independent evaluations formed without a press release attached.

By the time Zhipu confirmed the name, the model already had a week of unbiased usage behind it. That is a considerably stronger position than a launch post nobody trusts.

Cut Lead Review From Hours To Minutes

Sign up for a free trial of Attio, the agentic CRM.

Ask Attio to build a daily workflow that surfaces the deals that need your attention today, like anything with a stage change, a recent reply, or a new signal in the last 24 hours.

Review your pipeline in Claude, synced live from Attio via MCP.

That's it.

⚖️ The Sibling Problem

Now the contrast worth sitting with, because it says something about how open-weight announcements work.

GLM-5.3 launched on August 14, positioned as the strongest open-weights coding model currently available, claiming a 50 percent improvement over GLM-5.2 from post-training alone on an unchanged base. Zhipu said weights would follow in roughly two weeks pending security review, and did not disclose a license.

Two weeks later, those weights are still pending. The only way to use GLM-5.3 remains Zhipu's own paid coding plan.

GLM-5.3-Flash, announced twelve days later, shipped MIT-licensed weights to Hugging Face on the same day. MIT is the clean version, no revenue thresholds, no bespoke terms, no commercial gates. You can self-host it this week.

So within one model family you have both versions of what open weights currently means. One announced as open with the artifact arriving later under undisclosed terms. One announced and downloadable the same afternoon under the most permissive license there is.

That is not an accusation. A security review before releasing a model explicitly trained on vulnerability discovery is a defensible reason to wait. But it is a useful reminder that the announcement and the license are separate things, and the gap between them varies enormously even inside a single company.

What It Adds Up To

For anyone building, GLM-5.3-Flash is unusually easy to evaluate because there is nothing gating it. MIT weights are on Hugging Face now, the API is $0.15 per million input tokens, and if you were already testing Ox Alpha you have results.

Where it likely fits is the high-volume tier. An 18-billion active-parameter model with native multimodal input and a million-token context, at a tenth of its own flagship's price, is aimed squarely at classification, extraction, document processing, and agent loops where you are making thousands of calls and the frontier model is overkill. Approaching Opus 4.8 on Zhipu's internal coding benchmark is a vendor claim on a vendor benchmark, so treat that as a reason to test rather than a reason to migrate.

The broader signal is the release strategy. Zhipu has now twice shipped a major model anonymously, let real developers use it without knowing whose it was, and revealed it only after independent impressions formed. In a year when every benchmark chart is vendor-run and half the launch claims cannot be reproduced, letting the usage data arrive before the marketing is a genuinely smart move. Expect others to copy it.

Top 5 In AI Research 🔬

The stories moving fast beyond today's headlines:

🛠️ Tools That Are Hot Right Now!

  • 🤗 Hugging Face - where the MIT-licensed GLM-5.3-Flash weights landed on announcement day.

  • 🔀 OpenRouter - where Ox Alpha ran anonymously for six days, and the fastest way to compare it against what you use now.

  • 🧑‍💻 OpenCode - the open-source terminal agent GLM models plug into directly, no vendor lock-in.

  • 📊 Artificial Analysis - independent rankings, the correction to every internal benchmark in this issue.

What's The Recap?

A free model called Ox Alpha showed up on a routing platform on August 20 with no owner attached. Developers hammered it for six days. On August 26 Zhipu confirmed it was GLM-5.3-Flash, though the community had already worked out the name, since GLM-5.3 was text-only and Ox Alpha took images and video.

It runs 320 billion parameters with only 18 billion active, takes text, image, and video natively, handles a million-token context, and costs $0.15 per million input tokens. That is a tenth of the flagship rate, and MIT weights went up the same day.

Zhipu also did this in February with Pony Alpha, which turned out to be GLM-5. Ship anonymously, let people form opinions without a benchmark chart telling them what to think, then put your name on it.

The best detail is inside Zhipu's own lineup. Flagship GLM-5.3 was announced on August 14 as the strongest open-weights coding model and still has no published weights. The cheaper sibling shipped MIT weights on day one. Same company, same family, two very different meanings of open.

Login or Subscribe to participate

Stay building. 🤖

Recommended for you

View all
caret-right