ToolNavs Find Useful AI Tools
Submit Sign in

DeepSeek V4.1 Flash

With a 1M context and 384K of output, DeepSeek V4.1 Flash now absorbs traffic once handled by V4 Pro.

DeepSeek put the first landing point of its new architecture on the smallest model. V4.1 Flash handles native vision, a 1M context and 384K of output, with API concurrency of 2500. The company says it beats V4 Pro on performance, cost, speed and total time, and from September 14 V4 Pro requests route to it. The line proves native vision and throughput on a small model first.