通义千问发布 Qwen3.8-Omni-Flash:1M 上下文的全模态模型,主打智能体式音视频理解和工具调用Qwen Releases Qwen3.8-Omni-Flash: a 1M-Context Omni Model Built for Agentic Audio-Video Understanding and Tool Calling
阿里的通义千问团队发了 Qwen3.8-Omni-Flash,说是自家第一个围绕智能体能力搭起来的全模态模型。它能吃进文本、图像、音频和视频,吐出来的还是文本。音视频理解、推理和工具调用全放在同一个模型里。官方把它的做事顺序说得很简单:先看懂内容,再规划任务,然后调用工具执行,最后交付结果。
Alibaba's Qwen team released Qwen3.8-Omni-Flash, which it calls its first omni model built around agentic capability. It takes in text, images, audio and video, and still emits text. Audio-video understanding, reasoning and tool calling all sit inside a single model. The team describes its workflow simply: understand the content first, then plan the task, then call tools to execute, and finally deliver the result.

文本、图像、音频、视频都能吃进去,吐出来的是文本。音视频理解、推理和工具调用全塞在同一个模型里。官方给的做事顺序就四步:看懂内容、规划任务、拿工具执行、交出结果。发布时只开 API,没放权重。
一句话概览
阿里的通义千问团队发了 Qwen3.8-Omni-Flash,说是自家第一个围绕智能体能力搭起来的全模态模型。它能吃进文本、图像、音频和视频,吐出来的还是文本。音视频理解、推理和工具调用全放在同一个模型里。官方把它的做事顺序说得很简单:先看懂内容,再规划任务,然后调用工具执行,最后交付结果。
能落地吗?今天是以托管 API 的形式可用。它已经上到通义云、阿里云百炼和 Qwen Studio。发布时没有放出权重,所以想自己部署暂时没门。
Qwen3.8-Omni-Flash 是个什么
它站在 Qwen3.8-Flash-Next 这个架构上,那个基座模型是 2026 年 8 月带着开放权重发的。
上下文窗口 1M tokens。通义云上写的是最大输入 991K、最大输出 131K,最长推理 262K tokens。
输出只有文本。百炼文档里提醒开发者:要生成语音的话用 Qwen3.5-Omni。思考默认开着,reasoning_effort 是 xhigh;设成 none 就把思考关掉。
API 同时吃 DashScope 和 OpenAI 两套协议,Chat Completions 和 Responses API 都能用。函数调用、联网搜索、结构化输出、上下文缓存、批量调用都支持。
长视频上,它换了个看法
多数视频模型拿到一个长文件,都是从头到尾扫一遍——哪怕答案其实只在其中 3 分钟的镜头里。
通义研究团队描述的是另一条路:从问题出发,让智能体自己决定该看哪段、听哪段,再分几轮由粗到细地收集证据。算力和 tokens 都花在真正要紧的片段上。
团队公布的 OmniVideoBench 结果是:准确率从 63.4 涨到 67.8,token 用量从 145,736 降到 79,117,少了大约 45.7%。
官方给的成绩单
下面这些数字都出自通义自己。到发稿为止,还没有看到独立复现的结果。
- 29 项评测的平均分,比 Qwen3.5-Omni-Plus 高出 25% 以上
- WildClawBench-MM 涨了 36.5 分,AgenticVBench 涨了 22.3 分
- UniClawBench 拿到 69.6
- LongAudioSpan 涨 8.3 分,OmniVideoBench 涨 9.6 分
- OmniCap-IF 的 CSR 和 ISR 分别涨 8.5 分和 14.1 分
研究团队的说法是,音视频表现已经贴近 Gemini 3.8 Flash,整体音频表现则超过它。官方那条推文把智能体能力的提升总结成一句话:WildClawBench-MM 和 UniClawBench 上平均 +19.5 分。
价格和几条硬限制
通义云的价格是每百万输入 tokens 0.15 美元、每百万输出 tokens 0.47 美元,隐式缓存命中按每百万 tokens 0.016 美元算。
跟 Qwen3.5-Omni-Plus 比,团队报的降本幅度很大:音频输入每小时便宜 98% 以上,音视频输入每小时便宜 93% 以上。那条推文给的视频输入降幅是大约 89%。
百炼文档里几条值得记住的限制:
- 视频用链接传,最长 2 小时、最大 2GB
- 音频最长 3 小时
- 音频输入覆盖 113 种语言和方言
- 视频采到 15 fps 时结果稳定
- 用 use_multichannel 可以支持双声道立体声和四声道 FOA 空间音频
- 6 个地域可用:北京、新加坡、中国香港、东京、法兰克福、弗吉尼亚
用 OpenAI SDK 调起来就几行:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DASHSCOPE_API_KEY"],
base_url=os.environ["DASHSCOPE_BASE_URL"],
)
completion = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[{"role": "user", "content": [
{"type": "video_url", "video_url": {"url": os.environ["VIDEO_URL"]}},
{"type": "text", "text": "List the key moments with timestamps."},
]}],
modalities=["text"],
stream=True,
)
for chunk in completion:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
配套开源:Qwen-MM-Plugins 和 Qwen-Live Harness
模型只吐文本,所以处理媒体的活儿得交给工具。通义团队为此开源了两个项目。
Qwen-MM-Plugins 已经以 Apache-2.0 上线,口号是"让任何智能体外壳天生就是多模态的"。每一项能力都打包成一个 Skill,外加一个可选的 MCP server。它的引导安装器支持 Claude Code、CodeBuddy、Codex、Qoder、OpenClaw、Qwen Code 和 Gemini CLI。
Omni 这一系列能力,对应发布时的几个演示:
- omni-memory:给一段长视频建一份音视频记忆
- omni-video2note:把教程视频变成带插图的 PDF
- omni-chatcut:音乐转 MV、影视解说、保住说话人音色的视频翻译
还有个核心插件,让主模型能直接读本地的图片和视频帧。README 也坦白了一个当前缺口:多数外壳还没法把音频原生喂给主模型,音频这会儿还得绕 API。
关键要点
- Qwen3.8-Omni-Flash 吃文本、图像、音频、视频,吐文本
- 1M tokens 上下文,支持函数调用、联网搜索,思考默认开着
- 换上智能体式感知后,OmniVideoBench 从 63.4 提到 67.8,token 少用约 45.7%
- 通义云价格:每百万 tokens 输入 0.15 美元、输出 0.47 美元
- 发布时只有 API,另有 Apache-2.0 的 Qwen-MM-Plugins 给各种外壳接
本文为原文的完整中文翻译,按整句语义用中文习惯重写,配图取自原文,另补入公开媒体报道与项目仓库的公开截图。原文作者 Asif Razzaq,2026 年 9 月 18 日发布。基准成绩与价格均按原文口径翻译,厂商自报数据已在正文中标注,未作补充或删减。
It takes in text, images, audio and video, and emits text. Audio-video understanding, reasoning and tool calling all live in one model. The official workflow is four steps: understand the content, plan the task, execute with tools, deliver the result. At launch it's API-only, with no weights released.
The One-Line Summary
Alibaba's Qwen team released Qwen3.8-Omni-Flash, which it calls its first omni model built around agentic capability. It takes in text, images, audio and video, and still emits text. Audio-video understanding, reasoning and tool calling all sit inside a single model. The team describes its workflow simply: understand the content first, then plan the task, then call tools to execute, and finally deliver the result.
Will it ship? Today it's available as a hosted API. It's already on Qwen Cloud, Alibaba Cloud Model Studio and Qwen Studio. No weights were released at launch, so self-hosting isn't an option for now.
What Qwen3.8-Omni-Flash Is
It builds on the Qwen3.8-Flash-Next architecture, a base model released with open weights in August 2026.

The context window is 1M tokens. Qwen Cloud lists a maximum input of 991K, a maximum output of 131K, and a maximum reasoning length of 262K tokens.
Output is text only. The Model Studio docs remind developers: use Qwen3.5-Omni if you need speech generation. Thinking is on by default with reasoning_effort set to xhigh; setting it to none turns thinking off.
The API speaks both DashScope and OpenAI protocols, and works with Chat Completions and the Responses API. Function calling, web search, structured output, context caching and batch calls are all supported.
On Long Video, It Changed Its Approach
Most video models take a long file and scan it front to back — even when the answer only lives in three minutes of footage somewhere inside.
The Qwen research team describes a different route: start from the question, let the agent decide which segments to watch and which to listen to, then gather evidence over several passes from coarse to fine. Compute and tokens go to the segments that actually matter.
The team's published OmniVideoBench result: accuracy rises from 63.4 to 67.8, while token usage drops from 145,736 to 79,117 — about 45.7% less.
The Scorecard From the Team
All the numbers below come from Qwen itself. As of publication, no independent reproduction has been seen.
- Across 29 benchmarks, the average score is more than 25% higher than Qwen3.5-Omni-Plus
- WildClawBench-MM is up 36.5 points, and AgenticVBench is up 22.3 points
- UniClawBench scores 69.6
- LongAudioSpan is up 8.3 points, and OmniVideoBench is up 9.6 points
- On OmniCap-IF, CSR and ISR are up 8.5 points and 14.1 points respectively
The research team says audio-video performance is now close to Gemini 3.8 Flash, and overall audio performance surpasses it. The official tweet sums up the agentic gains in one line: an average of +19.5 points on WildClawBench-MM and UniClawBench.
Pricing and a Few Hard Limits
Qwen Cloud pricing is $0.15 per million input tokens and $0.47 per million output tokens, with implicit cache hits billed at $0.016 per million tokens.
Compared with Qwen3.5-Omni-Plus, the team reports steep cost reductions: audio input is more than 98% cheaper per hour, and audio-video input is more than 93% cheaper per hour. The tweet puts video input at roughly 89% lower.
A few limits from the Model Studio docs worth remembering:
- Video is passed as a link, up to 2 hours long and up to 2GB
- Audio can run up to 3 hours
- Audio input covers 113 languages and dialects
- Results are stable when video is sampled at 15 fps
- use_multichannel supports two-channel stereo and four-channel FOA spatial audio
- Available in 6 regions: Beijing, Singapore, Hong Kong (China), Tokyo, Frankfurt and Virginia
Calling it with the OpenAI SDK takes just a few lines:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DASHSCOPE_API_KEY"],
base_url=os.environ["DASHSCOPE_BASE_URL"],
)
completion = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[{"role": "user", "content": [
{"type": "video_url", "video_url": {"url": os.environ["VIDEO_URL"]}},
{"type": "text", "text": "List the key moments with timestamps."},
]}],
modalities=["text"],
stream=True,
)
for chunk in completion:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
Companion Open Source: Qwen-MM-Plugins and Qwen-Live Harness
The model only emits text, so handling media has to be handed off to tools. The Qwen team open-sourced two projects for that.
Qwen-MM-Plugins is already live under Apache-2.0, with the tagline "make any agent harness multimodal by default." Each capability is packaged as a Skill plus an optional MCP server. Its guided installer supports Claude Code, CodeBuddy, Codex, Qoder, OpenClaw, Qwen Code and Gemini CLI.
The Omni line of capabilities maps to several demos from the launch:
- omni-memory: build an audio-video memory for a long video
- omni-video2note: turn a tutorial video into an illustrated PDF
- omni-chatcut: music-to-MV, film and TV commentary, and video translation that preserves the speaker's voice
There's also a core plugin that lets the main model read local images and video frames directly. The README also admits a current gap: most harnesses still can't feed audio natively to the main model, so audio has to go around the API for now.
Key Takeaways
- Qwen3.8-Omni-Flash takes text, images, audio and video, and emits text
- 1M-token context, with function calling and web search support, and thinking on by default
- With agentic perception, OmniVideoBench rises from 63.4 to 67.8 while using about 45.7% fewer tokens
- Qwen Cloud pricing: $0.15 per million input tokens, $0.47 per million output tokens
- At launch it's API-only, plus Apache-2.0 Qwen-MM-Plugins to hook into various harnesses
This is a complete translation of the original article, rewritten sentence by sentence into natural English, with images taken from the original and supplemented by public screenshots from media coverage and project repositories. Original author: Asif Razzaq, published September 18, 2026. Benchmark scores and pricing are translated as stated in the original, vendor-reported figures are flagged in the body text, and nothing has been added or removed.