The mystery surrounding Ox Alpha has been solved. The model was not a long-running anonymous project from an unknown lab, but a short-term stealth preview of a GLM-series model from Z.ai, also known as Zhipu AI. It appeared on OpenRouter as stealth/ox-alpha on August 20, 2026. Six days later, Z.ai confirmed its official name: GLM-5.3-Flash.
2
5
10
What made Ox Alpha a sensation was not simply the secrecy. During the preview, developers could use a free reasoning model aimed at coding and agentic workflows. Early community tests also suggested that it was unusually capable at real software-engineering tasks.
What was Ox Alpha?
OpenRouter described Ox Alpha as a reasoning model for coding, sustained agentic work and production workloads. Its intended uses included long-horizon software engineering, complex reasoning and workflows that combine text with visual context.
30
39
The preview’s headline specifications were:
- Model ID:
stealth/ox-alpha
- Listed: August 20, 2026
- Context window: 1,048,576 tokens—roughly 1 million tokens
- Maximum output: 131,072 tokens
- Input modalities: Text, images and video
- Workflow features: Tool or function calling and structured outputs
- Price during the preview: Free for both input and output
6
12
14
Those specifications made Ox Alpha attractive for large codebases, long-running conversations, video analysis and multi-step automation. But a large context window and a zero-dollar price tag do not, by themselves, establish that a model is better than commercial alternatives. The real draw was its early performance on software-engineering tasks.
Why did people call it “Niu Lai”?
“Ox” translates naturally to the Chinese character 牛 (niú), meaning ox or cattle. Chinese developers therefore gave the model the more memorable nickname “Niu Lai” (牛来).
The name also played on a viral Chinese animated film titled Niu Lai, roughly translated as Here Comes the Ox. The film’s title had become an online meme around the time Ox Alpha appeared, helping the nickname spread.
15
“Niu Lai” was a community nickname, not a name announced by Z.ai. It also did not indicate any product, ownership or copyright relationship between the model and the film.
14
15
What does the early 8/10 DeepSWE result show?
The first widely circulated performance result came from a small community test. Ox Alpha reportedly solved eight of 10 real-world software-engineering tasks on DeepSWE, producing an 80% Pass@1 score. In the same early comparison, Claude Fable 5 scored 65% and GPT-5.6 Sol scored 52%.
1
3
48
| Model |
Early 10-task DeepSWE result |
| Ox Alpha |
8/10 (80%) |
| Claude Fable 5 |
65% |
| GPT-5.6 Sol |
52% |
DeepSWE is designed to test whether an agent can understand a real codebase, identify a problem and submit a working fix—not merely answer an isolated programming question.
49
57 That gives the result more practical relevance than a conventional code-generation demo.
Still, the number came from just 10 tasks reported by the community. Task selection, prompts, runtime settings and evaluation procedures can all affect the outcome. The careful conclusion is therefore that Ox Alpha performed impressively in an early 10-task run, not that it definitively defeated every competing model.
Later reports said Ox Alpha solved 63 of the full 113-task DeepSWE evaluation and came close to Claude Opus 4.8.
5 A full-suite result would be more informative than a 10-task sample, but public reporting did not provide enough detail about the configuration and reproducibility of that evaluation. It should therefore be treated as a reported result rather than an independently verified standard conclusion.
How did it reach the top of OpenRouter?
Ox Alpha combined several ingredients that are almost tailor-made for viral adoption: an unknown identity, free access, a roughly million-token context window, image and video input, and a clear focus on coding and agentic tasks.
3
6
10
The model quickly rose to the top of OpenRouter’s usage rankings. Bloomberg-related reporting described it as one of the largest launches in the marketplace’s history and said its usage exceeded DeepSeek’s by more than twofold. Other reports said it briefly ended DeepSeek’s long run at the top of a related developer-platform ranking and set a single-day usage record.
5
10
16
These figures should not be casually combined. OpenRouter rankings, OpenCode rankings, session counts and token volume are different measurements; none should automatically be interpreted as a user count or a model-quality ranking. What is clear is that the free window and the model’s apparent capabilities generated an unusually intense developer response.
Stripe CEO Patrick Collison also described Ox Alpha as “very impressive” on social media.
59 His comment helps illustrate the level of industry attention, but it was not an independent evaluation of the model’s performance.
How did developers identify Zhipu before the reveal?
Before the official announcement, developers used black-box testing to connect Ox Alpha to Zhipu. No single clue proved the case, but several technical signals pointed in the same direction.
1. A consistent 75-token tokenizer offset
Researchers compared Ox Alpha’s token counts with those of GLM-5.3 across multiple prompt sets. They reported that the underlying counts matched closely, with Ox Alpha showing a fixed difference of about 75 tokens.
1
17
When the same offset appears across different languages, code samples and emoji tests, one plausible explanation is an additional system prompt or request wrapper added by OpenRouter or the provider—not a completely different underlying vocabulary. A larger community analysis also reported exact matches across 95 test groups.
18
This kind of evidence can strongly suggest a shared tokenizer or closely related serving layer. It cannot, on its own, prove that two models use identical weights.
2. Matching video-token behavior
Another important clue came from video input. Community tests reportedly found that Ox Alpha and Zhipu’s GLM-5V-Turbo consumed identical numbers of tokens across multiple videos. They also appeared to share similar frame-sampling and resolution-scaling behavior.
17
24
Video-encoder behavior is harder to imitate through prompt wording than ordinary writing style, which made this clue especially interesting to observers. Even so, it remained an external black-box inference rather than a substitute for a formal disclosure by the developer.
3. Similar API errors and output behavior
Developers also reported similarities between Ox Alpha and GLM services in API error messages, responses to reasoning settings and general output behavior.
2
4 Taken together with the tokenizer and video-processing evidence, these serving-layer clues made the theory of a Zhipu pre-release model highly credible.
The widely repeated claim of “99% certainty” was an individual observer’s assessment of the evidence chain, not a statistically validated probability.
27 Guarded comments from people associated with Google DeepMind also fueled speculation about Gemini, but those hints were never confirmation. The uncertainty ended on August 26, when Z.ai confirmed that Ox Alpha belonged to the GLM family and named it GLM-5.3-Flash.
5
10
Ox Alpha was part of a wider “Stealth Alpha” pattern
Ox Alpha was not the first anonymous Alpha-branded preview to appear on OpenRouter. Earlier reported examples included:
- Pony Alpha → Zhipu GLM-5
- Hunter Alpha → Xiaomi MiMo-V2-Pro
- Elephant Alpha → Ant Group’s Inclusion AI
- Owl Alpha → Meituan LongCat-2.0
3
8
33
The launches followed a similar format: a model appears under an anonymous or semi-anonymous codename, often with free access, receives real-world developer traffic for a short period, and is later connected to its formal brand or model family.
34
38
The strategy has several plausible benefits:
- Reducing brand and origin bias. Users can experience the model before a company name or country association shapes their expectations.
- Collecting real-world feedback. Developers place the model inside codebases, tool chains and automation workflows that closed laboratory evaluations may not capture.
- Encouraging organic adoption. “Free, anonymous and apparently powerful” is a potent combination for generating experimentation and discussion.
- Testing infrastructure. A large free-usage window can stress-test throughput, long-context handling and reliability under agentic workloads.
These are reasonable explanations of the strategy, not motivations that every model company has publicly confirmed. In Ox Alpha’s case, the anonymous preview clearly generated substantial usage, technical investigation and international developer attention before Z.ai connected it to the GLM brand.
8
9
10
The bigger story is the launch strategy, not the secrecy
Ox Alpha’s story has two distinct phases. First came the anonymous OpenRouter preview: free, multimodal, long-context and unknown. Then came the product reveal, when Z.ai identified it as GLM-5.3-Flash.
2
10
The early 80% DeepSWE result indicates strong competitiveness on one set of real software-engineering tasks. But both the 10-task sample and later reports about the complete evaluation need to be interpreted alongside their testing configurations and verification status.
The more consequential development may be the emergence of anonymous previews as public testing grounds for model training, infrastructure and developer growth. For developers, the practical lesson remains straightforward: test a model against your own repositories, tools and data constraints rather than relying on a single leaderboard position or viral score.
Ox Alpha’s reveal also demonstrated the value of community fingerprinting. At the same time, it underscored an important distinction: even when black-box evidence is compelling, a high-probability inference is not the same thing as official confirmation.