ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
DeepSeek officially released V4.1 Flash: Native Vision, with the new architecture starting from small models

DeepSeek officially released V4.1 Flash: Native Vision, with the new architecture starting from small models

AI information Admin 14 views

DeepSeek officially released DeepSeek-V4.1-Flash. Officially defined as the smallest model in the new architecture series, natively supporting multimodal visual understanding, with core goals of higher capability ceiling, faster inference, higher throughput, and scaling to larger models. Flash is no longer just a "cheap version" but is becoming the first place for next-generation architectures.

The new architecture is first implemented in small models

In official benchmarks, V4.1 Flash achieved GPQA Diamond 90.9, Codeforces 3471, Terminal-Bench 2.1 90.6, and DeepSWE v1.1 74.2. Vision-related BabyVision (w/tools) scored 89.6, and Chartography (w/tools) scored 78.9.

These results come from DeepSeek's official tests and cannot be directly equated with independent third-party validation. But the direction is clear: DeepSeek is trying to squeeze inference, code, agents, and vision capabilities into a single high-throughput model, rather than continuing to split multiple dedicated entry points.

Flash began to take over Pro's tasks

DeepSeek V4.1 Flash has now launched its API, and model names have been changed to deepseek-flash. The previous generation V4 Flash and V4 Flash Vision Exp have been retired; the old model names remain compatible for the time being, and the routing is unified to V4.1 Flash.

More importantly, the Pro line. DeepSeek stated that after extensive testing, V4.1 Flash surpasses V4 Pro in performance, cost, speed, and total time. Starting at 12:00 on September 14, before the V4.1 Pro release, deepseek-v4-pro requests will be handled by V4.1 Flash and charged at Flash prices.

The cost of inference continues to decline

API specifications also reflect the priority of AI inference efficiency: V4.1 Flash supports 1M context, up to 384K output, and has an API concurrency limit of 2500; For comparison, the current V4 Pro has a concurrency limit of 500.

During off-peak periods, the input price per million tokens cached is $0.003, cache missed is $0.15, output is $0.6, and the price doubles during peak periods. Capability is moving closer to Pro, and prices remain at the Flash tier—these are the most direct commercial signals for this upgrade.

V4.1 Flash is more like a pioneer version of DeepSeek's new architecture: first using small models to verify native vision, high throughput, and inference efficiency, then scaling to larger models. The key next is how much V4.1 Pro will push the upper limit of this architecture's capabilities.

Recommended Tools

More