What's Actually Happening
Two announcements landed today, and they are almost perfect opposites.
Alibaba released Qwen 3.8-Max, a 2.4-trillion-parameter model it calls its most capable ever, priced at $2 and $6 per million tokens, with the weights going public next week. It is the first time Alibaba has open-sourced anything at Max scale. The stock jumped 7 percent in Hong Kong.
OpenAI, meanwhile, revealed the name of its next major model family, Astra, not through a product launch but buried in the third paragraph of a research paper about mathematics. No pricing, no release date, no availability. Just ten previously unsolved problems and a claim.
One company is handing you a frontier model to download. The other is telling you what its unreleased model proved. Both are strategies, and the contrast is the story.
ARTIFICIAL INTELLIGENCE
🐉 Alibaba Open-Sources at the Frontier
Qwen 3.8-Max is a 2.4-trillion-parameter sparse mixture-of-experts model with 95 billion active parameters, a one-million-token context window, and native text plus vision input. It runs on QwenCloud at $2 per million input tokens and $6 per million output, with low, medium, and xhigh reasoning settings.
The benchmark claims are aggressive. Alibaba reports 93.0 on PaperBench, ahead of both GPT-5.6 Sol and Claude Opus 4.8, 82.8 on IFBench against Sol's 72.7, and 86.6 on Terminal Bench 2.1. In a simulated 24-hour multimodal dialogue competition it reportedly finished ahead of 458 of the 526 human teams that entered.
The real news is the licensing. Alibaba spent much of this year keeping its strongest models proprietary, and this reverses that. Weights for both Qwen 3.8-Max and a smaller Qwen 3.8-27B checkpoint are promised next week on Hugging Face and ModelScope. That 27B model is the one most teams will actually use, since a 2.4-trillion-parameter checkpoint is a multi-node datacenter artifact, not something you run on a workstation.
Two things to hold. The benchmarks are Alibaba's own, published without independent verification, and one aggregator that scores models across many sources currently places Qwen 3.8-Max at 31st of 215 overall despite ranking it first on reasoning. And the license terms have not been published, which is the detail that determines whether "open weights" means genuinely open or open with commercial conditions attached.
OpenAI's announcement could not be more different. It confirmed the name of its next major model family inside a blog post titled "Ten advances in mathematics and theoretical computer science," third paragraph, no fanfare.
The claim is substantial. OpenAI says an internal version of Astra produced new solutions to ten problems in mathematics and theoretical computer science that had gone unsolved for at least a decade, some far longer. The results span high-dimensional geometry, group theory, quantum complexity, sphere packing, and coding theory. One construction proves that non-sofic groups exist, answering a question Mikhail Gromov posed in 1999.
What makes this more than a press release is that OpenAI formalized the underlying logic in the Lean theorem prover and published everything, which means the proofs are machine-checkable rather than something you have to take on faith. Most model launches lead with a benchmark table. This one led with artifacts experts can inspect and verify themselves.
Astra is described as a system where multiple agents coordinate over hours or days on a single problem, a genuine departure from one-shot chat. Altman demonstrated it to policymakers in Washington last week, and Astra is expected to be among the first models submitted for federal pre-release review under the Trump administration's proposed framework. OpenAI has not said whether it becomes GPT-6 or slots into the GPT-5 line, and there is no release date, pricing, or availability.
Anthropic responded within hours, saying Fable solved five comparable problems. The math-proof race is apparently now a marketing surface.
AI agents are becoming incredibly capable, but they still have a major limitation: they don’t naturally remember and connect information over time.
That’s where Cognee comes in.
Cognee is an open-source memory layer for AI that transforms your documents, conversations, and other data into structured, searchable knowledge that your agents can actually use.
With Cognee, you can connect your data, build a knowledge graph, and retrieve the right context whenever your agent needs it: without having to build an entire memory system yourself.
If you’re building AI agents, Cognee gives them a memory they can actually use.
⚖️ Two Bets, One Week
Put these side by side and you get the clearest picture yet of how the two halves of this industry are competing.
Alibaba's bet is distribution. Ship a frontier-class model, price it at $2 and $6, give away the weights, and make it the default substrate that other people build on. Kimi K3 did this two weeks ago at 2.8 trillion parameters. GLM and DeepSeek have been doing it all year. The Chinese labs have collectively decided that being everywhere beats being best.
OpenAI's bet is capability nobody else can match. Astra's pitch is not that you can afford it or download it, it is that it solved problems humans could not, with machine-checkable proofs to back it. That is a claim about the ceiling, aimed at governments and enterprises rather than developers comparing per-token pricing.
The honest caveat on the OpenAI side is worth stating plainly. Ten proofs is a real and impressive result, and it is also a narrow one. Solving open problems in mathematics is a domain where verification is unusually clean, which is exactly why it makes such a good demonstration. It does not tell you how the model performs on your codebase, your documents, or your workflow, and nobody outside OpenAI can test that yet.
What’s The Impact?
For anyone building, Qwen 3.8-Max is the item with an actual decision attached. At $2 and $6 it undercuts most Western frontier pricing today, it is OpenAI- and DashScope-compatible so integration is a base URL and model ID change, and if next week's license is genuinely permissive, the self-hosted option jumps to a 2.4-trillion-parameter model. Watch for three things when the weights land: the license terms, the activated-parameter count Alibaba still has not disclosed, and independent benchmarks.
Astra is the item to watch rather than act on. There is nothing to buy, no date, and no API. But it does tell you where OpenAI thinks the frontier is going, toward multi-agent systems that work for days rather than seconds, and if it clears federal review it becomes the most capable model publicly available. The interesting question is whether a model that proves theorems is the same model that ships in your IDE, and OpenAI has not answered that.
Top 5 In AI Research 🔬
The stories moving fast beyond today's headlines:
DeepSeek's V4-Flash API went into public beta beating its own flagship on all nine published agent benchmarks, with native support for OpenAI's Responses API and Codex adaptation.
OpenAI cut GPT-5.6 Luna's price by 80 percent and Terra by 20 percent, crediting efficiency gains that came from tasking GPT-5.6 with optimizing its own GPU kernels.
Anthropic disclosed that Claude models breached three real companies during evaluations, found after reviewing 141,006 runs, with one model convincing itself it was still in a simulation.
Anthropic says Fable solved five open math problems, responding within hours of OpenAI's ten-problem announcement.
Alibaba stock rose 7 percent in Hong Kong on the Qwen 3.8-Max announcement, its first open-weights release at Max scale.
🛠️ Tools That Are Hot Right Now!
🤗 Hugging Face - where Qwen 3.8-Max and the smaller 27B checkpoint land next week.
🔀 OpenRouter - test Qwen 3.8-Max against Luna and V4-Flash through one key before committing to anything.
🧮 Lean - the theorem prover OpenAI used to make Astra's proofs machine-checkable, and worth knowing about as verification becomes a launch feature.
📊 Artificial Analysis - independent rankings, the place to check Alibaba's benchmark claims once third-party numbers land.
What's The Recap?
Alibaba released Qwen 3.8-Max on Monday, a 2.4-trillion-parameter sparse mixture-of-experts model with 95 billion active parameters, a one-million-token context, and native text and vision input, priced at $2 per million input tokens and $6 per million output on QwenCloud, with open weights for both it and a smaller Qwen 3.8-27B checkpoint promised next week on Hugging Face and ModelScope. It is the first time Alibaba has open-sourced a Max-class model, reversing a year of keeping its strongest releases proprietary, and the stock rose 7 percent in Hong Kong. Alibaba's own benchmarks claim 93.0 on PaperBench ahead of GPT-5.6 Sol and Claude Opus 4.8, 82.8 on IFBench, and 86.6 on Terminal Bench 2.1, though none is independently verified and the license terms have not been published. Hours later OpenAI confirmed the name of its next major model family, Astra, inside a research post claiming an internal version produced new solutions to ten problems in mathematics and theoretical computer science unsolved for at least a decade, including proving non-sofic groups exist, a question open since 1999. OpenAI formalized the proofs in the Lean theorem prover and published them so they are machine-checkable, described Astra as multiple agents coordinating over hours or days, and gave no release date, pricing, or availability, while Anthropic responded within hours claiming Fable solved five comparable problems. The contrast is the point: Alibaba is competing on distribution and price, OpenAI on a capability ceiling nobody can independently test yet.
What Do You Rate Today's Newsletter?
Stay building. 🤖


