Z.AI Serves GLM-5.3 Flash on 100,000 Chinese-Made Chips, Reducing Reliance on Nvidia
Key Takeaways
- •Z.ai said GLM-5.3 Flash is handling online queries using 100,000 China-made chips.
- •The company released the 320-billion-parameter model under the MIT open source license on August 26 after previewing it on August 20.
- •GLM-5.3 Flash is priced at $0.075 per million input tokens and $0.25 per million output tokens until Sept. 9, before rising afterward.
- •The model became the most downloaded model on OpenRouter after attracting many developers in a blind test.
- •The development comes amid tighter U.S. export controls and Chinese government moves to curb purchases of Nvidia AI chips.

Chinese AI vendor Z.ai said it used 100,000 China-made chips to handle online queries to its latest AI model, GLM-5.3 Flash. The move shows how Chinese vendors are reducing their dependence on U.S. chipmaker Nvidia, while also highlighting the trend toward greater optimization of AI models.
Z.ai, formerly known as Zhipu AI, is a Beijing-based startup that grew out of Tsinghua University's research community and is counted among China's leading AI model developers. Its new model is open-weight and multimodal, meaning its trained parameters can be downloaded and run by third parties, and it is low-cost, with 320 billion parameters. Z.ai originally previewed the model under the code name Ox Alpha on August 20 and officially released it under the MIT open source license — one of the most permissive open source licenses, permitting commercial use, modification and redistribution — on August 26. The model is specialized for long-context processing, vision-driven agentic tasks and code synthesis.
GLM-5.3 Flash costs $0.075 per million input tokens and $0.25 per million output tokens from now until Sept. 9. After that, the price will be $0.15 per million input tokens and $0.50 per million output tokens. Tokens — the chunks of text, roughly a word or word fragment, that models read and generate — are the standard billing unit across AI services. Comparatively, GPT-5.6 Luna from OpenAI costs $0.20 per million input tokens and $2 per million output tokens, while Anthropic's Claude Opus 5 is $5 per million input tokens and $5 per million output tokens.
The Geopolitical Backdrop
The Chinese vendor's decision to use purely Chinese chip providers comes as China aims to rely less on Nvidia. China's pivot toward self-reliance follows years in which the U.S. tightened export controls — a regime that began with October 2022 restrictions on advanced accelerators such as Nvidia's A100 and H100 — to prevent Chinese tech vendors from obtaining the most powerful Nvidia chips. Although some of these controls have eased, in September 2025 China's government ordered tech giants such as Alibaba and ByteDance to stop buying Nvidia AI chips, and in November Beijing banned all foreign AI chips from state-funded data centers.
The Advancement of Chinese Hardware
Despite these moves, many observers still saw China as behind in the infrastructure contest. Clusters of that size have been the signature of the largest U.S. AI labs — xAI's Colossus supercomputer in Memphis, for example, was reported to run on roughly 100,000 GPUs. However, Z.ai's strategy suggests that the gap between China and the U.S. in AI chips may be narrowing, especially with the increased focus on inference over the past year amid the sharp rise of agentic AI. Inference — running a trained model to answer queries, as opposed to training, the far more compute-intensive process of building a model — is precisely the job Z.ai's 100,000 chips are doing for GLM-5.3 Flash.
“For training the largest frontier models, Nvidia still has a significant advantage in chip performance, networking and the broader software ecosystem,” said Kashyap Kompella, CEO and founder of RPA2AI Research. “But inference is a different story.”
He added that Z.ai's ability to serve GLM-5.3 Flash at scale using Chinese chips indicates that Chinese hardware is already good enough for many high-volume inference workloads.
“This distinction is important because inference is likely to become the larger AI chip market over time,” Kompella said. “Chinese companies may not need exact chip-for-chip parity with Nvidia if they can compensate through efficient model architectures, software optimization and large clusters of domestically made AI chips.”
However, the first signals of Chinese vendors' move toward independence began with DeepSeek — the Hangzhou-based lab whose low-cost models drew global attention in early 2025 — rewriting Nvidia CUDA drivers to make inference faster, said Bradley Shimmin, an analyst at Futurum Group. He said that while Chinese vendors are responding to constraints first imposed by the U.S. government and later by the Chinese government, another factor is that companies “need to do more with the watt hours they have available to them.”
Vendors such as Google, with its TPU series, and Amazon, with its Trainium chips, are focusing on vertical integration, in which the AI stack is optimized to use AI models, infrastructure and software together for the most efficient AI processing.
“What we're witnessing here is this real realization that AI value is driven by economics of the stack, as much as it is about the capabilities of the models,” Shimmin said. “What these guys are doing is simply showcasing the value of a little bit of investment in optimizing that hardware.”
The Model
As for GLM-5.3 Flash itself, the fact that it received strong developer mindshare in a blind test is significant, Kompella said. Many developers flocked to the model early this month and used it without knowing it came from a Chinese vendor; it quickly became the most downloaded model on OpenRouter, an API service that lets developers route requests across models from many different providers.
“That is more interesting than another benchmark claiming that a Chinese model is close to a U.S. frontier model,” Kompella said, adding that it shows that Chinese open-weight models “continue to function as a counterweight to higher industry pricing.”
For enterprises, this means they need to design their AI system “so workloads can be routed across models based on capability, quality, cost and strategic considerations rather than becoming dependent on a single model provider,” he continued.
Although the model's immediate attraction of a large number of developers is a good sign for Z.ai, the vendor still faces challenges ahead. For one, it must prove that the chips it is using can continue to deliver reliability, economy and scale, Kompella said.
“At the absolute frontier of model training, access to compute also remains a meaningful constraint because Nvidia's advantage is much harder to replicate there,” he said.
The vendor will also need to convert U.S. and European enterprises that remain cautious about using Chinese-hosted AI services due to data security, regulatory and procurement concerns, Kompella continued.