Zhipu Reveals Ox Alpha as GLM 5.3 Flash Running Entirely on Chinese AI Chips

AI

The mystery surrounding Ox Alpha is officially over. Chinese AI company Z.ai has confirmed that the anonymous model released through OpenRouter and OpenCode was an early version of GLM 5.3 Flash, its new 320 billion parameter multimodal model. More significantly for China's semiconductor industry, Z.ai says every request generated during the anonymous test was served entirely through domestically developed AI accelerators, demonstrating that a frontier class model can now be deployed at significant scale without relying on NVIDIA GPUs.

Ox Alpha appeared anonymously on August 20 and quickly attracted attention among developers because of its coding performance, multimodal capabilities and extremely generous access. The model supported a context window reaching 1 million tokens alongside text, image and video inputs. OpenCode promoted the test with near unlimited free usage and claimed infrastructure capacity reaching 100 trillion tokens per day. That figure represents advertised serving capacity rather than independently verified daily traffic, but the scale of the trial remains notable because Z.ai has now confirmed that the entire workload was handled through Chinese AI hardware.

"Before release, we tested GLM 5.3 Flash anonymously as ox alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week, with all of this traffic served on Chinese AI chips."
— Quote by: Z.ai

Z.ai has not identified the accelerator manufacturers used for the deployment. This means speculation connecting the infrastructure directly to Huawei Ascend or another specific Chinese chip family remains unconfirmed. What the company does disclose is considerably more useful technically. GLM 5.3 Flash was deployed across tens of thousands of domestically developed accelerators connected through high bandwidth interconnects, with Z.ai building a specialized inference engine on top of SGLang to overcome limitations in individual chip compute performance, memory capacity and memory bandwidth.

The serving architecture uses separate Encode, Prefill and Decode worker pools, allowing multimodal processing, prompt computation and token generation to scale independently across the cluster. Z.ai also implemented W8A8 quantization, mixed INT8, FP8 and BF16 cache quantization, tensor parallelism, ReplaySSM and Layer Split techniques to reduce memory pressure and increase hardware utilization. The company says these optimizations produced a 3x improvement in complete serving performance compared with its original implementation on the same hardware, eventually reaching hardware efficiency and per token cost comparable with mainstream NVIDIA GPUs.

GLM 5.3 Flash itself was designed around similar efficiency goals. The model contains 320 billion total parameters but activates only 18 billion during inference, compared with 32 billion active parameters in the earlier GLM 4.5 generation. It also reduces the number of model layers from 92 to 45 and combines linear attention with sparse attention, allowing the architecture to preserve long context capabilities without processing every token through a conventional full attention mechanism.

Z.ai says these changes reduce attention computation by approximately 3x and shrink KV cache requirements by 4.4x compared with GLM 5.3. An additional technology called IndexPool compresses indexer key vectors when operating with context windows reaching 1 million tokens, further reducing memory overhead. The model was pretrained on a 30 trillion token multimodal corpus and uses Manifold Constrained Hyper Connections to improve scaling efficiency between layers.

Performance is competitive despite the lower compute requirements. GLM 5.3 Flash scored 63.4 on DeepSWE v1.1 compared with 46.2 for GLM 5.2 and 58.0 for Claude Opus 4.8. On AutomationBench it reached 48.8 compared with 26.2 for GLM 5.2 and 41.0 for Claude Opus 4.8. Z.ai also reports a score of 57 on Artificial Analysis Intelligence Index v4.1.1 at a discounted cost of only 0.045$ per task. As with all vendor published benchmark comparisons, real performance will vary considerably according to workload, agent framework, context length and serving configuration.

The hardware achievement may ultimately be more important than the benchmark scores. China's AI industry has spent years trying to reduce its reliance on NVIDIA as export restrictions complicate access to leading American accelerators. Domestic hardware has historically faced disadvantages in individual accelerator performance, memory bandwidth and software maturity, but Z.ai demonstrates how architecture and software optimization can compensate for some of those limitations at cluster scale. This follows the wider shift explored in China's domestic AI chip market potentially capturing nearly 90% of high end shipments, as developers increasingly optimize models directly around locally available silicon.

The result also reinforces a point previously raised around Huawei narrowing the gap with NVIDIA through large scale accelerator deployments. Chinese accelerators do not necessarily need to defeat Blackwell or Vera Rubin individually to become strategically viable. If thousands of domestic chips can be connected efficiently, supplied reliably and supported by an increasingly mature software stack, Chinese AI companies gain another path toward scaling frontier models without depending entirely on CUDA based infrastructure.

The 100 trillion token headline is impressive, but the more important achievement is that Z.ai publicly demonstrated frontier model serving across tens of thousands of Chinese accelerators and claims it brought the economics close to mainstream NVIDIA hardware after a 3x software optimization improvement.

That is where the AI hardware competition is changing. NVIDIA still holds an enormous advantage through CUDA, networking, HBM, developer tools and raw accelerator performance, but China increasingly appears willing to compensate through scale and aggressive hardware aware software development. GLM 5.3 Flash is particularly well suited to that strategy because its 18 billion active parameters, hybrid attention architecture and reduced KV cache requirements lower the pressure placed on weaker individual accelerators.

Z.ai has not proven that Chinese GPUs have reached Blackwell level performance. It has demonstrated something potentially more important for China's AI independence: domestic hardware can already serve a competitive frontier model at meaningful production scale.

Is matching NVIDIA efficiency through software and massive domestic clusters enough for China, or does it still need individual AI accelerators capable of competing directly with Blackwell and Vera Rubin?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

Fujitsu MONAKA Stacks 2 nm CPU Cores Over 5 nm SRAM as MONAKA X Targets 1.4 nm in 2029

Next
Next

Where Winds Meet Expands Its 2.0 Era With High Stakes Extraction Mode