ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Qwen 3.8-Omni-Flash released: Full-modal AI begins directly executing audio and video tasks

Qwen 3.8-Omni-Flash released: Full-modal AI begins directly executing audio and video tasks

AI information Admin 6 views

Qwen released Qwen 3.8-Omni-Flash, defining it as the first full-modal model built around agent capabilities. It not only handles text, images, audio, and video, but also integrates understanding, reasoning, planning, and tool calls into a single workflow, aiming to advance audio and video AI from "understanding content" to "completing tasks."

From understanding video to executing tasks

In the past, full-modal models focused mainly on recognition, Q&A, and summarization, but Qwen 3.8-Omni-Flash shifts the focus to execution. Official scenarios include automatic video logging, short video translation, movie summaries generation, and continuous tool calls to perform multi-step tasks around long videos.

Alibaba Cloud documentation shows that the model supports Function Calling, Web Search, and the default inference mode, with input covering text, images, audio, and video, and output as text. For developers, this means audio and video content can directly serve as the Agent's task context, rather than as separate material attachments.

1 million tokens is not just "packed in more"

Qwen3.8-Omni-Flash provides 1 million token context. A more critical change is "proxy viewing": for long videos, the model can first perform coarse-grained searches and then actively identify the segments that need to be carefully examined, rather than processing the entire segment from start to finish.

Qwen's published OmniVideoBench data shows that this method improves accuracy while significantly reducing token consumption. The official report also states that it has made significant improvements over the previous generation in WildClawBench-MM, UniClawBench, and other proxy tests, with audio and video capabilities approaching Gemini 3.8 Flash, though these results are still mainly based on manufacturer tests.

Costs are declining, and plugins are beginning to fill the ecosystem

Price may have a greater impact on implementation than the rankings. Qwen stated that compared to Qwen 3.5-Omni-Plus, Qwen 3.8-Omni-Flash significantly reduces video input costs, making it easier for conference analysis, long-form video retrieval, and persistent audio-video agents to move from demos to high-frequency calls.

The open-source Qwen-MM-Plugins now cover video-to-note, long-form video memory, video editing, and Omni Skill Creator, and support Agent Harnesses such as Claude Code, Codex, Gemini CLI, OpenClaw, and Qwen Code. Qwen-Live Harness continues to target real-time audio-video interaction and task orchestration.

What Qwen3.8-Omni-Flash truly deserves attention is not "another stronger multimodal model," but its capability to turn audio-video understanding into sustainable Agent calls. The next round of competition will likely shift from who can see more accurately and who can complete longer, more complex real-world tasks at lower cost.

Recommended Tools

More