ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
GLM-5.3-FlashX launch: up to 200 tokens/s, Zhipu pushes large model competition into the Infra

GLM-5.3-FlashX launch: up to 200 tokens/s, Zhipu pushes large model competition into the Infra

AI information Admin 1 views

Zhipu officially launched GLM-5.3-FlashX, with a maximum inference speed of 200 tokens/s, API simultaneous opening, and Model Key GLM-5.3-FlashX. Compared to the existing GLM-5.3-Flash, the new version is about 5 times faster, with a price increase of 2.5 times. The competition among AI models has shifted from "who is smarter" to "who can output stably faster and cheaper."

From Ox Alpha to FlashX

GLM-5.3-Flash previously anonymously logged into OpenCode and OpenRouter as Ox Alpha . Z.ai revealed that within a week of launch, the model became one of the most frequently accessed models on the two major platforms, processing over 62 trillion tokens in six days, with high concurrency demands directly pushing pressure onto inference infrastructure.

Compared to retraining a larger model, FlashX is more like a production "performance layering": maintaining the Flash series' capability positioning while increasing output throughput to 200 tokens/s, enabling Coding Agents, real-time conversations, and tool calls to achieve shorter wait times.

100,000 domestic chips support faster inference

Speed is not just about increasing machines. Zhipu states that GLM-5.3-Flash production inference runs on a cluster of over 100,000 domestically produced AI accelerator chips, and the inference system is rebuilt for memory, bandwidth, and new model architectures.

Official optimizations include W8A8 quantization, mixed-precision cache, Layer Split, and Encode-Prefill-Decode decoupling architecture. After multiple rounds of optimization, end-to-end serving performance on the same hardware is about three times higher than the initial version.

The large model competition has entered the infrastructure stage

FlashX makes speed a higher-priced API tier, essentially pricing it for "low latency." For enterprises, every step an agent takes must wait for the model to return, and tokens/s directly determine whether the task chain can be compressed from tens of seconds to an interactionable range.

As the gap in model capability narrows, inference engines, chip adaptation, cluster scheduling, and cost per unit token will become new moats. The significance of GLM-5.3-FlashX is not just a 200 tokens/s metric, but more so that domestic computing power is beginning to undertake large-scale real-world calls and directly translate Infra optimization into product differentiation.

Recommended Tools

More