APAIIF 亞太人工智能產業總會APAIIFAI Knowledge
Latest AI Technology

Zhipu AI GLM-5.3-Flash Revealed: Mystery 'Ox Alpha' Model Unmasked, 320B MoE Architecture Sets Record 70 Million Daily Requests

August 30, 20265 Views
Zhipu AI GLM-5.3-Flash Revealed: Mystery 'Ox Alpha' Model Unmasked, 320B MoE Architecture Sets Record 70 Million Daily Requests
智譜AI
GLM-5.3-Flash
開源AI模型
混合專家架構
中國AI

Zhipu AI GLM-5.3-Flash Revealed: Mystery 'Ox Alpha' Model Unmasked

On August 26, 2026, Chinese AI company Zhipu AI (Z.ai) officially confirmed to Bloomberg that the mystery model anonymously tested on OpenRouter as "Ox Alpha" for one week was its latest flagship model, GLM-5.3-Flash. This revelation ended a week of speculation in the AI community and exposed the full picture of a new AI model launch strategy.

The Mysterious Appearance of Ox Alpha

On August 20, 2026, an anonymous model named stealth/ox-alpha quietly appeared on AI model routing platform OpenRouter. The model's characteristics were remarkable:

  • Completely Free: No payment required
  • Massive Context Window: Supporting 1,048,576 tokens (approximately 1 million tokens)
  • Strong Agentic Coding Capabilities: Quickly became the preferred tool for AI agents like Claude Code and Hermes Agent

News spread rapidly through the developer community. Early testing showed the model claimed to score 80% on the DeepSWE benchmark, significantly outperforming competitors at the time. However, subsequent more comprehensive community testing placed its performance in the "mid-pack of frontier models" rather than at the absolute top.

Explosive Usage Growth

Ox Alpha's free strategy achieved remarkable results. By August 24, 2026, the model's daily request volume had surpassed 70 million, making it one of the highest-usage models on the OpenRouter platform.

Behind this figure lies the strong demand from AI developers for high-performance, low-cost inference services. However, during the anonymous testing period, the model's prompts and responses were retained by the operator. While stated not to be used for training, the lack of a formal data processing agreement prompted security analysts to warn against using sensitive or production-grade data with this model.

GLM-5.3-Flash Technical Specifications

After the official reveal, Zhipu AI published the complete technical specifications for GLM-5.3-Flash:

Specification Details
Architecture Mixture-of-Experts (MoE)
Total Parameters 320B
Active Parameters 18B
Context Window 1,048,576 tokens (~1M)
Multimodal Support Text, image, and video input
Open Source License MIT License
Release Platform Hugging Face (zai-org/GLM-5.3-Flash)
Pricing $0.15/M input tokens, $0.50/M output tokens

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, supporting mixed input of text, images, and video, significantly expanding its application scenarios.

A Milestone for Domestic AI Chips

Zhipu AI specifically emphasized that during the Ox Alpha anonymous testing phase, the model ran entirely on domestic Chinese AI chips. This statement carries significant strategic importance.

Against the backdrop of strict U.S. export restrictions on AI chips to China, Zhipu AI's ability to support 70 million daily requests with domestic chips demonstrates the actual capabilities of China's AI infrastructure. This is not only a technical achievement but also a powerful response to external doubts about China's AI computing capacity.

The GLM-5.3 Family: Two Models with Different Positioning

It's worth noting that GLM-5.3-Flash and the simultaneously released GLM-5.3 are two models with distinctly different positioning:

GLM-5.3 (Released August 14, 2026):

  • 743B/744B parameter scale
  • Focused on cyber defense and security tasks
  • CyberGym benchmark score 84.5%, AutomationBench score 48.2%
  • Staged, controlled open-source strategy (not fully open as of end of August)

GLM-5.3-Flash (Released August 26, 2026):

  • 320B total / 18B active parameter MoE architecture
  • Targeting broad agentic application scenarios
  • Immediately fully open-sourced under MIT license
  • Lower inference costs, suitable for large-scale deployment

Lessons from the "Stealth Launch" Strategy

Ox Alpha's launch method represents an emerging AI model promotion strategy: collecting large-scale evaluation data in real production environments anonymously and for free, while building developer mindshare, then officially launching a paid version after accumulating sufficient data.

The effectiveness of this strategy is evident: before the official launch, GLM-5.3-Flash had already established widespread recognition and a usage base in the global developer community.

Impact on the Asia-Pacific AI Ecosystem

The open-source release of GLM-5.3-Flash has significant implications for the Asia-Pacific AI ecosystem. According to CSIS analysis, multiple Asia-Pacific nations are actively pursuing "sovereign AI" strategies, seeking to reduce dependence on U.S. AI technology.

GLM-5.3-Flash's MIT license open-source strategy enables Asia-Pacific enterprises and research institutions to:

  • Deploy high-performance AI models locally without relying on overseas cloud services
  • Customize and fine-tune according to local needs
  • Reduce the overall cost of AI applications

Conclusion

The revelation of Zhipu AI GLM-5.3-Flash is not just a technology launch event — it's a concentrated demonstration of China's AI industry capabilities. From high-concurrency service supported by domestic chips, to full open-source under MIT license, to record-breaking usage volumes, GLM-5.3-Flash is reshaping the competitive landscape of global open-source AI models. For Asia-Pacific enterprises seeking AI technology diversification, this is an important development worth closely monitoring.

FAQ

Related Articles