ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Ant Group Releases Ling-3.1-flash: 560B Parameters Built for Agents and Long Tasks

Ant Group Releases Ling-3.1-flash: 560B Parameters Built for Agents and Long Tasks

AI information • Admin • • 6 views

On September 30, 2026, Ant Group's InclusionAI (Ling) announced its next-generation language model, Ling-3.1-flash, via its official WeChat account: 560 billion total parameters with only about 25 billion activated per inference, aimed at agent tasks, search, office software, and specialist applications. It is the first time the flash lineup has pushed parameter scale into the 500-billion class.

Nearly five times the parameters, and "flash" still means efficiency

The previous generation, Ling-3.0-flash, took the "small and fast" route: 124 billion total parameters and 5.1 billion activated, running at 606 tokens per second on a single request across four Blackwell cards thanks to its hybrid attention architecture, released under the MIT license and self-hosted by many developers. Ling-3.1-flash raises the bar to 560 billion total and 25 billion activated — nearly five times the scale of its predecessor.

Activated parameters make up roughly one-twentieth of the total — exactly the logic of the mixture-of-experts (MoE) architecture: the model stores a large set of "expert" modules and calls only the most relevant ones for each inference, decoupling knowledge capacity from inference cost. For the mechanics behind "huge parameters, modest activation," see this explainer: What is Mixture-of-Experts (MoE)? Why do so many popular models have huge parameter counts but modest activation.

One million tokens is the goal; the trial starts at 256K

According to the official announcement, Ling-3.1-flash is designed for a context window of up to one million tokens, targeting "real-world long tasks." But the current two-week free trial caps context at 256,000 tokens; after the trial, InclusionAI plans to open the larger window and release the model as open source.

This cadence follows Ling's established playbook: free trial first to gather real-scenario feedback, then open source. Given that several previous Ling models were released under the MIT license, the "open source after trial" promise carries real credibility.

The priority order reveals intent: from chatting to getting work done

The most telling detail in the announcement is the priority order — agent tasks come first, followed by search, office software, and specialist applications. That means Ling-3.1-flash is not optimized for general chat, but designed for long-chain work like "completing a job": multi-round tool calls, long-document processing, and cross-application collaboration — precisely where context length and stability matter most.

The long-task track: what Ling is chasing

The shift from "small and fast" to "large and complete" mirrors where the whole industry has moved: in 2026 the contest is no longer just about chat quality, but about which model can reliably finish a real task. In agent scenarios, a model must hold its goal steady across dozens of interaction rounds, call tools correctly, and digest extremely long inputs — exactly the problem the large-parameters-plus-long-context combination is meant to solve.

For developers, the most practical move is to take advantage of the two-week free trial: 256,000 tokens of context is enough to validate most long-task workflows, and if the open-source release after the trial keeps the MIT license, enterprises get a 560-billion-parameter-class long-task model they can self-host at zero licensing cost — a rarity among Chinese-built models. For the previous generation's measured performance, see: Ling 3.0 Flash decoding speeds up: 606 tok/s on a single request with four Blackwell cards.

Recommended Tools

More