Skip to content
AI ConnectPowered by VELENTIS
AI-generated2 min

Zhipu AI Releases Open-Weights Model GLM-5.3-Flash Following Viral Test Run

Zhipu AI has launched its open-weights model GLM-5.3-Flash under an MIT license after the system handled over 62 trillion tokens during an anonymous test run under the codename Ox Alpha.

This article was AI-generated and published automatically. Context, labelling and all sources at the end of the article.

(KI-generiertes Symbolbild: Gemini / AI Connect)

Chinese AI company Zhipu AI officially unveiled its new open-weights model GLM-5.3-Flash on August 26, 2026, releasing the weights under a permissive MIT license. The launch follows weeks of speculation across the global developer community regarding a high-performance, mystery model that had appeared on benchmarking platforms under the pseudonym Ox Alpha. With this release, Zhipu AI makes a system specifically optimized for high throughput and cost-effective deployment freely accessible to developers and enterprises.

Architecturally, GLM-5.3-Flash employs a Mixture-of-Experts (MoE) design with 320 billion total parameters. However, only 18 billion parameters are activated per individual inference step, significantly curbing the computational footprint per generated token. At the same time, the model supports a long context window of up to one million tokens. To manage the immense key-value cache overhead typically associated with large context lengths, Zhipu AI integrated a hybrid approach combining sparse and linear attention mechanisms.

Prior to its formal unmasking, the model generated substantial momentum within the AI scene. Running on inference platforms such as OpenRouter and OpenCode under the codename Ox Alpha, the system processed more than 62 trillion tokens in anonymous production settings. Many developers praised its rapid latency and solid reasoning capabilities without knowing the underlying vendor or architectural specifications.

A key element of the release centers on the computing infrastructure used to power the deployment. According to Zhipu AI, the entire inference traffic during the testing period was handled by a massive cluster comprising 100,000 domestically manufactured Chinese AI accelerator chips. The reliable operation of this cluster highlights ongoing advances in China's hardware ecosystem to support production-scale inference independently of Western supply chains.

By opting for an MIT license, Zhipu AI grants users wide latitude for commercial deployment, fine-tuning, and on-premises integration. The blend of reduced serving costs, parameter efficiency, and full open-weights availability is poised to intensify competition among open foundation models designed for long-context workloads.

What this means for you

For engineering teams and enterprises, GLM-5.3-Flash delivers a permissive, cost-efficient open-weights option for heavy workloads requiring massive context windows. It also demonstrates that large-scale MoE serving can be operated efficiently across alternative hardware clusters.

Perspectives

Coverage: 3× Other

One story, several angles: how each source frames the topic, each with a verbatim quote.

  • emergent.shOther

    The source focuses on the model's speed, multimodal architecture, and permissive MIT open-source licensing as a developer-friendly alternative to proprietary systems.

    Original quote

    Zhipu AI has officially released GLM-5.3-Flash, a multimodal artificial intelligence model designed for speed and accessibility.

    emergent.sh
  • scmp.comOther

    The report highlights the surge in Zhipu's shares and the viral stealth trial, emphasizing that the model ran on domestic Chinese chips as a milestone in reducing reliance on US hardware.

    Original quote

    Zhipu’s shares closed more than 12 per cent higher at HK$1,160 in Hong Kong on Thursday.

    scmp.com

Source classification is maintained editorially (political spectrum only where consensus is broad; vendor communication is PR, not journalism). Unlabelled sources are unclassified: we do not guess.

Evidence

Solidly sourced
69/100
  • Zhipu AI released GLM-5.3-Flash on August 26, 2026, under an MIT license.

    verified
  • Operating under the pseudonym Ox Alpha, the model processed over 62 trillion tokens across OpenRouter and OpenCode.

    single source
  • The entire inference traffic was served on a cluster of 100,000 domestically produced Chinese AI chips.

    single source

The evidence score is computed, not hand-set: from confidence, the number of sources and the share of verified statements.

Source & transparency

As of: August 28, 2026

AI-generatedAI-generated: produced automatically from vetted sources with technical quality checks (source, quote and figure verification); no human sign-off of each item before publication

Sources
3
Verified statements
1 / 3
Evidence score
69Solidly sourced

Want to put this into practice?

We connect you with suitable AI providers from the DACH region, free of charge and without obligation.

What's next?