News by Nicholas Cabel

AI desk

Z.ai's GLM-5.3-Flash is the open-weight model behind OpenRouter's Ox Alpha

AIAI summaryNicholas Cabel

Z.ai released GLM-5.3-Flash, a 320B MoE with 18B active under MIT, confirmed it was OpenRouter's mystery Ox Alpha, and priced it at about a tenth of GLM-5.2.

A hand lifting a white linen cloth off a small bronze bull statue
AI-generated illustration

Z.ai released GLM-5.3-Flash, a 320B MoE with 18B active under MIT, confirmed it was OpenRouter's mystery Ox Alpha, and priced it at about a tenth of GLM-5.2.

Key points

  • 320B total parameters with 18B active per token. The weights are on Hugging Face as zai-org/GLM-5.3-Flash under an MIT licence, so they are free to download and use commercially; on Z.ai's API the model code is glm-5.3-flash, per its model docs. Hugging Face already lists 99 quantised community builds at the time of writing.
  • Z.ai's price list: $0.15 per 1M input tokens, $0.03 for cached input, $0.50 per 1M output. GLM-5.2 and GLM-5.3 both sit at $1.40 in and $4.40 out, so the company's claim of a tenth of the price works out to about 11 percent at list rates. The model docs also give a discounted cost of $0.045 per task on Artificial Analysis's Intelligence Index.
  • 1M-token context with 128K output, per the model docs (SiliconANGLE gives the exact 131,072); inputs are text, images, video and files. Z.ai calls it the GLM-5 series' first natively multimodal model, which is its own framing: its vision-focused GLM-5V-Turbo, on OpenRouter since April 1, already took images and video.
  • Z.ai's comparison table sets it against GLM-5.2, DeepSeek-V4-Vision-Exp, Claude Opus 4.8, GPT-5.6 Terra and Gemini 3.7 Flash. On Z.ai's own runs it scores 84.3 on Terminal-Bench 2.1 and places second on AutomationBench (48.8 against GLM-5.2's 26.2). It tops the six on GDPval-AA v2 at 1773, a test the model card says Artificial Analysis ran; Artificial Analysis's live leaderboard, checked on Sept. 11, has it at 1669, behind entries for Claude Fable 5.1, Claude Opus 5 and Muse Spark 1.3.
  • The model card describes a freshly trained base model with hybrid sparse-plus-linear attention and mHC (manifold-constrained hyper-connections), pretrained on a 30T-token multimodal corpus, and lists serving support in SGLang, vLLM, TokenSpeed, Transformers, KTransformers and Unsloth. The model docs say attention compute and KV-cache size drop by 3.0x and 4.4x versus GLM-5.3.
  • OpenRouter hosted it free as Ox Alpha, with no developer named, for about a week before launch. Quartz reports that Z.ai confirmed on Aug. 26 that it was the maker, and OpenRouter now lists the model as z-ai/glm-5.3-flash. Z.ai's docs say all the Ox Alpha traffic ran on Chinese AI chips. Quartz reports that the company put the serving cluster at 100,000 chips, a figure CNBC was unable to verify; the docs themselves describe tens of thousands of accelerators developed in China.

“approaching Claude Opus 4.8 on coding and agentic benchmarks.” — Z.ai, GLM-5.3-Flash model card on Hugging Face

Why it matters

For a small shop, an MIT-licensed model that can take a whole repo plus screenshots and screen recordings in one 1M-token prompt, and bills fifty cents per million output tokens, changes the maths on agentic coding loops that would drain a budget on frontier APIs. One caution on self-hosting: the 18B active count buys throughput, not a smaller footprint, and 320B total weights still mean on the order of 160 GB at 4-bit by back-of-envelope arithmetic, so a local deployment is a multi-GPU node or an aggressive quant rather than a single-card box. Most of the headline scores are Z.ai's own runs, so wait for more independent testing before trusting the rankings, but the price-to-context ratio alone earns it a slot in a model router.

This is an AI-written summary of the reporting credited above and the other sources linked in the text, read and edited by Nicholas before publishing. The facts and any quote belong to those sources; the wording is ours. Read the original.