The Beijing-based company, known in China as Zhipu AI, officially released GLM-5.3-Flash on August 26 after quietly testing it online under the codename Ox Alpha. The anonymous model had appeared on OpenRouter and OpenCode, where it quickly attracted heavy developer usage before Z.AI revealed its identity.

Z.AI said the preview traffic was handled by a cluster containing around 100,000 Chinese-made AI chips. The company did not disclose which manufacturers supplied the hardware, and the claim has not been independently verified.

In an official post, Z.AI said GLM-5.3-Flash was served at scale on Chinese AI accelerators using an inference stack specifically optimized around the limitations of domestic hardware. The company says those optimizations delivered roughly a threefold improvement in serving performance compared with its initial baseline.

Ox Alpha turns out to be GLM-5.3-Flash

Before its formal release, GLM-5.3-Flash spent several days online as the mysterious Ox Alpha model. The stealth launch allowed Z.AI to expose the system to real-world developer workloads without initially revealing who had built it.

According to the company, Ox Alpha processed around 62 trillion tokens across OpenRouter and OpenCode before its identity was disclosed. OpenRouter data also showed unusually strong usage during the model's first days on the platform, particularly for coding workloads.

The approach is not entirely new among Chinese AI developers. Companies have increasingly used anonymous releases to test models under real-world conditions before attaching a brand name to them, reducing some of the expectations that normally accompany a major launch.

GLM-5.3-Flash follows Zhipu AI's GLM-5, which was released earlier this year as a large open-weight model focused heavily on coding and long-running agent tasks.

A 320-billion-parameter model with only 18 billion active

GLM-5.3-Flash contains 320 billion total parameters but activates around 18 billion for each request. The Mixture-of-Experts design is intended to reduce the amount of computation required during inference while preserving the capabilities of a much larger model.

It is also the first natively multimodal model in the GLM-5 family, allowing it to process text alongside images and video. Z.AI is positioning the system for coding, agentic workflows, visual understanding and other long-context tasks.

The company says it redesigned both the architecture and serving software around efficiency. Among other techniques, the system separates different stages of inference into independently managed compute pools and uses a mixture of quantization and memory optimizations to reduce pressure on hardware bandwidth.

That matters because Chinese AI accelerators generally remain constrained by lower memory bandwidth and computing performance compared with Nvidia's most advanced processors. Limited access to high-end chips has repeatedly been identified as one of the main obstacles facing China's AI industry.

China tests whether domestic chips can handle AI at scale

The significance of GLM-5.3-Flash goes beyond the model itself. If Z.AI's claims hold up, the deployment would show that Chinese-designed accelerators can support large-scale inference workloads for a competitive AI model without relying directly on Nvidia GPUs.

Z.AI has not named the specific chips used in the 100,000-unit cluster. The company has previously worked with domestic semiconductor companies including Huawei, Cambricon Technologies and Moore Threads. Cambricon and Moore Threads have both announced support for the new GLM model.

The distinction between inference and training is important, however. Serving an already trained model is generally less demanding than training a frontier-scale system from scratch. The deployment therefore does not establish that Chinese chips have reached performance parity with Nvidia across every stage of AI development.

Z.AI's own claims about hardware efficiency and cost have also not yet been independently validated. Still, the scale of the deployment makes GLM-5.3-Flash a notable example of China's effort to reduce its dependence on foreign AI hardware.

Low pricing targets developers

Z.AI is also competing aggressively on cost. GLM-5.3-Flash is priced far below the company's full GLM-5.3 model, with rates around $0.15 per million input tokens and $0.50 per million output tokens at standard pricing.

The combination of open weights, relatively low inference requirements and aggressive API pricing could make the model attractive to developers looking for alternatives to more expensive frontier systems from U.S. providers.

Early feedback has been mixed. Some developers have praised Ox Alpha for handling difficult coding and debugging tasks, and Stripe CEO Patrick Collison publicly described his experience with the model as impressive. Others have reported hallucinations, unfinished tasks and slower-than-expected code generation.

That leaves the model's long-term competitive position uncertain, but its launch illustrates an increasingly important trend: Chinese AI companies are no longer competing only on benchmark scores. They are also pushing on price, open-weight distribution and the ability to run large workloads on domestic infrastructure.

Z.AI shares rise after the reveal

Investors responded positively to the announcement. Z.AI shares rose more than 12% in Hong Kong trading on August 27 following the disclosure that Ox Alpha was GLM-5.3-Flash and that the model had been served using Chinese-made chips.

Z.AI was founded in Beijing in 2019 out of research connected to Tsinghua University and has since become one of China's most prominent large-language-model developers. Its GLM family competes with systems from Alibaba, Moonshot AI, DeepSeek and other Chinese laboratories.

For China, the bigger question is whether projects such as GLM-5.3-Flash can turn domestic AI hardware from a substitute used under export restrictions into a genuinely competitive computing ecosystem. Running a heavily used model on Chinese accelerators is one step toward that goal, but training performance, energy efficiency and independent validation will remain critical tests.