◆ THE LAB · Models
Z.ai ran its new model unbranded for six days
A model called ox-alpha ran anonymously and free on OpenRouter from August 20, judged on its output alone before anyone knew who built it. Z.ai claimed it as GLM-5.3-Flash on August 26 and published the weights under an MIT licence.
Chandni Melwani
Founder & Editor
Aug 28, 2026 · 2 MIN READ
The News
On August 20, 2026 a model listed only as stealth/ox-alpha appeared on OpenRouter, the marketplace developers route requests through, described as "a reasoning model designed for coding, sustained agentic work, and production workloads." It ran free and unattributed for six days. On August 26 Z.ai claimed it as GLM-5.3-Flash and published the weights on Hugging Face under an MIT licence. The model has 320B total parameters with 18B active per query — a sparse mixture-of-experts design, where only a fraction of the network fires on any given token — plus a hybrid sparse-and-linear attention architecture and native multimodality. Z.ai says it approaches Claude Opus 4.8 on coding and agentic benchmarks. Almost every figure in its comparison table is self-reported, against models and settings Z.ai chose; one GDPval row is marked as evaluated by Artificial Analysis. Artificial Analysis, which benchmarks models independently, scores it 57 on its Intelligence Index against a median of 29 for open-weight models of similar size.
Know More
- — The stealth listing went up on OpenRouter on August 20, 2026 and now carries the line "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash." The page states that prompts and completions were retained by the provider and not used for training.
- — Architecture: 320B total parameters, 18B active, hybrid sparse-plus-linear attention, natively multimodal across text and images, MIT licence, weights at zai-org/GLM-5.3-Flash.
- — Z.ai's own benchmark claims, self-reported: 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE. The lab chose the comparison models and the evaluation settings.
- — Artificial Analysis scores it 57 on its Intelligence Index, against a median of 29 among open-weight models of comparable size — and records it generating 150M output tokens on that evaluation where the median is 110M.
- — Speed depends on who is serving it, not on the model. Across hosts Artificial Analysis measures Databricks at 272.9 tokens per second and SiliconFlow at roughly 47, with Z.ai's own API at 49.
- — Pricing is promotional: $0.075 per million input tokens and $0.25 per million output. Z.ai's pricing page shows these as a 50% discount off a list price of $0.15 and $0.50, ending "at 24:00 on September 9, 2026 (UTC+8, Singapore time)."
For six days in August, a model with no name worth having did a lot of work. It went up on OpenRouter on August 20 as stealth/ox-alpha, free, described in one line as a reasoning model for coding and sustained agentic work, and developers pointed their agents at it without knowing whose it was. On August 26, Z.ai said it was theirs and published the weights as GLM-5.3-Flash under an MIT licence.
320B parameters in total with 18B firing per token is the kind of spec that gets a launch post written about it, and it is what sparse mixture-of-experts buys you — a big model with a small model’s serving bill — plus hybrid sparse-and-linear attention and native multimodality. Z.ai’s own comparison table puts it near Claude Opus 4.8 on coding and agentic work.
The six days are the more interesting object. Anonymous distribution on the venue Stripe is currently buying produced something a launch post cannot: real reaction formed before anyone could attribute it. Nobody was grading a Chinese lab, or a frontier lab, or a lab they had opinions about — provenance is usually the first thing an audience reaches for, and here it was unavailable. Free access muddies it, since price pulls traffic on its own. It is still the closest thing this beat has had to a blind taste test.
Everyone else grading this model had a stake in the grade. Z.ai reported its own benchmark figures, against rivals it chose. Artificial Analysis, scoring independently, broadly backs the capability claim and then adds two findings the marketing omits. Six days of unbranded use is the one piece of evidence here that nobody’s press office shaped.
So what: an MIT licence on a model this capable means the serving layer is where the margin now sits — the same weights run at 272.9 tokens per second on Databricks and about 47 on SiliconFlow, and that spread is somebody’s business.
Related
Room for Disagreement
Z.ai says GLM-5.3-Flash approaches Claude Opus 4.8 on coding and agentic benchmarks, and points at 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE. Those are the lab's own runs, against rivals and settings the lab picked, and that is the normal condition of a model launch rather than a scandal. Artificial Analysis, testing independently, largely agrees on capability — 57 on its Intelligence Index where the median for open-weight models of this size is 29 is a real result, not a rounding of the marketing. Z.ai does not measure the rest. The same independent pass records the model as slower than average and more verbose than average, generating 150M output tokens on an evaluation where the median is 110M. Verbosity bills tokens and adds latency on every step of an agentic loop. Both accounts are accurate, and only one of them prices the whole thing.
Frequently Asked Questions
Why does it matter that the model ran anonymously first?
Nothing in the listing said which lab built it, so for six days developers picked it or dropped it on the output alone. On a beat where almost every capability claim arrives attached to a logo — and where the logo does a lot of the persuading — that is a rare thing to have happened by accident. It is not a controlled experiment: the model was also free, which pulls in traffic on its own and makes usage a poor proxy for quality. Those developers reacted before they could attribute, and nobody can reconstruct that reaction after the fact.
How do Z.ai's numbers and the independent scores differ?
They measure different things and both are attributable. Z.ai's comparison table is self-reported, using the lab's own choice of rival models and settings, and says the model approaches Claude Opus 4.8 on coding and agentic work. Artificial Analysis runs its own evaluations and puts it at 57 on its Intelligence Index — genuinely strong for an open-weight model this size, against a peer median of 29 — while also finding it slower than average and more verbose than average, neither of which appears in Z.ai's material. Strong and slow are both true.
What should a developer budgeting against these prices watch for?
Two ways. The $0.075 and $0.25 per million tokens are a 50% launch discount that Z.ai's own pricing page says ends on September 9, 2026 — the list prices behind it are $0.15 and $0.50. And tokens-per-second figures belong to whichever host is serving the weights, not to the model: the same weights run at 272.9 tokens per second on Databricks and around 47 on SiliconFlow. An MIT licence means anyone can serve it, so both numbers move depending on where you run it.
Sources
- Z.ai — GLM-5.3-Flash model card and weights (MIT licence)
- OpenRouter — stealth/ox-alpha listing, created August 20, 2026, since revealed as GLM-5.3-Flash
- Artificial Analysis — GLM-5.3-Flash: Intelligence Index, output tokens, verbosity
- Artificial Analysis — GLM-5.3-Flash output speed by host, including Databricks and SiliconFlow
- Z.AI Developer Document — Pricing: GLM-5.3-Flash 50% discount to September 9, 2026
- SiliconANGLE — Z.ai open-sources Ox Alpha model as GLM-5.3-Flash (August 26, 2026)
Chandni Melwani
Chandni Melwani is the founder and editor of New in AI, covering AI agents, M&A, and enterprise adoption. She holds a Master's in Management of Artificial Intelligence from Queen's University and brings a practitioner's perspective from her work in Data and AI leadership.
Follow ↗The Lab Briefing
Get the five-beat briefing before the market opens.
One email. Frontier labs, AI at work, e-commerce, and the capital crossing borders — decoded daily.
No spam. Unsubscribe anytime.
More from The Lab
Safety
Anthropic's automated researcher cheated on 2.4% of runs
On August 28, Anthropic published research in which a Claude agent acted as an autonomous alignment researcher, proposing and testing fixes for alignment problems in another model. Across 1,601 monitored transcripts, the agent tried to cheat in 39 of them.
Chandni Melwani · Aug 30, 2026
Safety
OpenAI put a number on what safety costs: about a fifth of the compute it watches
The largest planned frontier training run is still on hold, and the monitoring around it runs at roughly 20% of the inference compute being monitored.
Chandni Melwani · Global · Aug 18, 2026
Labs
The White House says Moonshot AI trained its model on Anthropic's, and accessed restricted Nvidia chips in Thailand
Washington built export controls to keep advanced chips out of Chinese hands. The allegation is that a Chinese lab reached them in Thailand instead, and that the model capability never needed to cross a border at all.
Chandni Melwani · China · Jul 23, 2026