Z.ai confirmed on August 26, 2026, that the anonymous free preview called Ox Alpha was GLM 5.3 Flash—a 320B parameter, natively multimodal model with 18B active per token and MIT licensed weights. The model accepts text, images, and video, supports roughly one million tokens of context, and is priced at $0.15 per mi...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Z.ai, the Chinese AI company formerly known as Zhipu AI, reveal about the anonymous “Ox Alpha” model that appeared free and without. Article summary: Z.ai disclosed on August 26 that the previously anonymous, free “Ox Alpha” preview was **GLM-5.3-Flash**—the first natively multimodal GLM-5 model. It is an open-weight, low-cost coding and agent model whose stealth rele. Topic tags: general, general web, user generated, documentation. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks,
The mystery ended on August 26, 2026: Z.ai confirmed that Ox Alpha, the anonymous model offered through OpenRouter and OpenCode, was GLM-5.3-Flash. Z.ai presented it as the first natively multimodal model in the GLM-5 family, with open weights, a one-million-token context window, and a focus on coding and long-running agent tasks. 1540
The important story is broader than the model’s identity. Ox Alpha combined a free, anonymous preview with large-scale developer usage, then became a named and downloadable product. That approach let Z.ai test its infrastructure under real workloads while building interest around the eventual release.
GLM-5.3-Flash is a mixture-of-experts model with 320 billion total parameters and 18 billion active parameters per token. It natively accepts text, images, and video, while producing text output. The model’s listed context capacity is approximately one million tokens, although the exact displayed limit varies by platform: Z.ai documentation describes the one-million-token design, while OpenRouter lists a 1,310,720-token context window. 12340
Z.ai released the weights under the MIT License, making the model substantially more deployable than a conventional closed API-only system. The model was also listed under its public name on OpenRouter after the anonymous preview ended. 348
The release should not be confused with the larger GLM-5.3 product announced earlier in August. Reporting describes GLM-5.3-Flash as a separate, lower-cost 320B-A18B model, rather than the same SKU as the larger GLM-5.3 line. 10
Z.ai says GLM-5.3-Flash improves on GLM-5.2 while reducing serving costs. In the results cited around the launch, it scored 63.4 on DeepSWE v1.1, compared with 46.2 for GLM-5.2. Z.ai also reported a 29.0 score on its internal coding benchmark, close to the 29.5 reported for Claude Opus 4.8. 316
Those numbers are useful indicators of the company’s positioning, but they are not independent confirmation of parity with Anthropic’s model. The benchmark definitions, evaluation settings, prompting, and reproducibility details matter. The safest conclusion is that Z.ai is claiming strong coding and agent performance at a much lower cost—not that GLM-5.3-Flash has definitively surpassed closed frontier systems.
Ox Alpha’s most visible feature was that developers could try it without paying during the anonymous preview. That phase ended when the model received its public identity. Z.ai’s reported list price is $0.15 per million input tokens and $0.50 per million output tokens. 13
OpenRouter showed a lower promotional rate of $0.075 per million input tokens and $0.25 per million output tokens around the launch. 2 The distinction matters: the free preview, Z.ai’s list price, and a platform-specific launch discount are three different access conditions.
Even at the stated list price, the model’s economics are central to its appeal. Coding agents can consume large quantities of context and output, so lower token costs can influence model selection as much as a small benchmark advantage.
Z.ai said GLM-5.3-Flash was running entirely on Chinese-made AI accelerators. 15 Reporting on the serving system describes a customized SGLang-based stack using techniques including quantization, tensor parallelism, mixed cache formats, layer splitting, and a separate encode–prefill–decode architecture. The reported goal was to improve end-to-end serving performance while working within the constraints of domestic hardware. 25
That extends an earlier Z.ai narrative around the GLM family, which the company has said was trained on Huawei Ascend processors without Nvidia hardware. Reporting also notes that Z.ai’s inclusion on the U.S. Entity List restricts access to Nvidia’s H100, H200, and B200 accelerators. 26
The inference claim is strategically significant, but it should be read as a company and industry report rather than a fully audited comparison of hardware performance. The stronger verified takeaway is that Z.ai is presenting its model stack as capable of serving a large multimodal model without relying on Nvidia’s leading data-center GPUs.
The anonymous release appears to have followed a deliberate stealth-launch pattern: publish a capable model without attribution, offer it free through popular coding routes, observe how developers use it, and reveal the product after the system has accumulated real-world feedback. Z.ai has been associated with an earlier “Pony Alpha” test using a similar idea, while community investigators tried to identify Ox Alpha through tokenizer behavior, video-processing clues, and API-error patterns. 71113
Some release-period reports also cited unusually large usage totals and developer enthusiasm, but the available material does not independently verify every figure or quotation. Those claims are best treated as launch reporting, not audited adoption data. 14
GLM-5.3-Flash does not establish that Z.ai has overtaken OpenAI or Anthropic across the full range of frontier-model capabilities. It does demonstrate a more disruptive combination: multimodal input, very long context, open MIT-licensed weights, coding-agent specialization, and low API pricing in one release. 1540
For developers, that combination can lower switching costs and make self-hosting or multi-provider deployments more practical. For closed-model providers, it adds pressure on price and margins—especially in coding workloads where token volume, latency, deployment control, and infrastructure availability may matter as much as leaderboard position.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Z.ai confirmed on August 26, 2026, that the anonymous free preview called Ox Alpha was GLM 5.3 Flash—a 320B parameter, natively multimodal model with 18B active per token and MIT licensed weights.
Z.ai confirmed on August 26, 2026, that the anonymous free preview called Ox Alpha was GLM 5.3 Flash—a 320B parameter, natively multimodal model with 18B active per token and MIT licensed weights. The model accepts text, images, and video, supports roughly one million tokens of context, and is priced at $0.15 per million input tokens and $0.50 per million output tokens on Z.ai’s stated list price.
The stealth rollout gave Z.ai production scale coding agent traffic before the company attached its name, making the launch both a real world serving test and a high impact marketing strategy.