ToolNavs AI Tool Directory
Submit Sign in
Back to AI information
Zhipu open-source GLM-5.3-Flash, the 320B model has pushed AI prices to new lows

Zhipu open-source GLM-5.3-Flash, the 320B model has pushed AI prices to new lows

AI information Admin 4 views

Zhipu (Z.ai) officially launched and open-sourced the GLM-5.3-Flash (320B-A18B). It is the first native multimodal model in the GLM-5 series, supporting 1M token contexts and bringing API prices down to just a fraction of flagship models. Rather than simply pursuing bigger specs, this release feels more like a "capability/cost ratio" offensive.

320B parameters, only 18B activated

GLM-5.3-Flash has a total parameter count of 320B, with only 18B activated per inference, reducing computational load while retaining frontier model capabilities. Native multimodal and 1M context allow code, images, and long tasks to enter the same Agent workflow.

In the Artificial Analysis Intelligence Index, this model scored 57, ranking in the same tier as Claude Opus 4.8. Zhipu's self-developed Z.ai Code Bench also delivers similar programming performance, though the latter is based on vendor testing standards.

Ox Alpha reveals his identity

Before its official release, the model was anonymously previewed under the name Ox Alpha on OpenCode and OpenRouter. Zhipu claims that this traffic is entirely driven by Chinese AI chips; Now, model weights are open under the MIT license, allowing developers to deploy them themselves.

From anonymous stress testing to open weights, this approach has extended the validation scenarios of domestic computing power beyond model training to large-scale online inference.

Price is more aggressive than specs

Official API is 50% off for two weeks: $0.075 input per million tokens, $0.25 output, and $0.015 cache input per million tokens. The discount also covers third-party model aggregators, directly reducing the costs of AI programming and high-frequency agent calls.

According to Zhipa standards, GLM-5.3-Flash is typically priced at about one-tenth of GLM-5.3, with promotional periods at about one-twentieth, and the overall cost can be as low as one-forty-fifth that of Claude Opus 4.8. Currently, the model has entered ZCode, Coding Plan, Chat, and API.

The highlight of GLM-5.3-Flash is not just "Flash," but that cutting-edge capabilities, 1M context, open-source weighting, and low-cost APIs are all integrated into one model. If this cost structure can be sustained, model competition will shift more quickly from "who is strongest" to "who can scale strong capabilities."

Recommended Tools

More