Qwen 发布原生全模态模型 Qwen3.8-Omni-Flash,主打音视频智能体任务交付
Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.
Qwen 发布下一代原生全模态模型 Qwen3.8-Omni-Flash,支持文本、图像、音频和视频输入及 1M token 上下文窗口,29 项评测平均分较 Qwen3.5-Omni-Plus 提升超过 25%,音频输入每小时价格下降超过 98%,音视频输入每小时价格下降超过 93%。
官方发布同时开放 API 和开源插件与运行时,读者可以据此评估音视频智能体在剪辑、会议、实时交互等工作流中的落地方式。

QWEN-LIVE HARNESS
QWEN-MM-PLUGINS
QWEN3.8-OMNI-FLASH API
QWEN3.8-OMNI-FLASH-REALTIME API
Introduction#
Today, we are launching Qwen3.8-Omni-Flash, our next-generation native omnimodal model. Its core objective is to strengthen agent capabilities in real-world productivity scenarios, advancing omnimodal models from “understanding omnimodal content” to “planning tasks, calling tools, and completing creative work.” Building on general agentic capabilities in coding, text-based knowledge work, and GUI operation, Qwen3.8-Omni-Flash further extends agentic applications centered on audio and video, delivering strong results across workflows such as video editing, music video creation, film production and commentary, audio-visual summarization, and real-time conversations.

Figure 1. Qwen3.8-Omni-Flash and its applications in production.
- Qwen3.8-Omni-Flash — now available on the Qianwen AI Platform:
- Text, image, audio, and video inputs with a 1M-token context window.
Qwen3.8-Omni-Flash supports a 1M-token context window while maintaining text performance comparable to a text-only model of the same size and delivering significant improvements in omnimodal capabilities. Across 29 evaluations1, its average score improves by more than 25% over Qwen3.5-Omni-Plus; the API price per hour of audio input decreases by more than 98%, and the price per hour of audio-visual input decreases by more than 93%2. For audio-visual agents, coding, and long-horizon tasks, the model improves by 36.5 points on WildClawBench-MM and 22.3 points on AgenticVBench, while scoring a strong 69.6 on UniClawBench. Its core capabilities also improve significantly in long-form audio and audio-visual understanding, audio-visual reasoning, audio-visual captioning, and multi-speaker recognition. For example, it gains 8.3 points on LongAudioSpan and 9.6 points on OmniVideoBench; its OmniCap-IF CSR and ISR improve by 8.5 and 14.1 points, respectively; and its AliMeeting DER and cpWER decrease from 88.11 / 89.61 to 3.35 / 17.18. By scaling data, context, and agentic environments, Qwen3.8-Omni-Flash achieves audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash. These advances also mean that audio and video are evolving from perceptual inputs into core media through which agents understand their environment, reason, and execute tasks.
1. The scope includes audio reasoning benchmarks: AliMeeting-test, AISHELL-4, MagicData-RAMC, MLC-SLM (en), WenetSpeech (Net | Meeting), FLEURS-60 ASR, FLEURS-60 S2TT, SpotSoundBench, MMAU, MMAR, MMSU, MuchoMusic-RUL, HumMusQA, MusTBench, Audio-MultiChallenge, WildSpeech, and VoiceBench; audio-visual reasoning benchmarks: DailyOmni, WorldSense, AVUT, JointAVBench, OmniCloze, OmniCap-IF, QIVD, OmniVideoBench, and StreamingBench; and audio-visual agent benchmarks: WildClawBench-MM, UniClawBench, and OmniGAIA.
2. Pricing methodology: hourly audio or audio-visual input prices are estimated as 30 times the input cost of two minutes of source material; audio-visual input uses 720p at 1 fps. Gemini 3.8 Flash uses media_resolution=high, Seed 2.0 Lite uses max_frame_tokens=384, and all other API parameters use their default values. Text input and output prices are in CNY per 1M tokens; Gemini and Muse prices are converted from USD at an exchange rate of 1 USD = 6.7191 CNY.

Audio and video are important media for bringing agents into real-world productivity scenarios, but they also introduce a new set of system-level challenges. Long-form audio and video are costly to store, transmit, and process across multiple rounds of inference; existing agent harness frameworks lack native support for these modalities; and workflows that connect omnimodal understanding with end-to-end task execution are still at an early stage. Addressing these challenges requires models, harness tools, and runtime environments to evolve together.
To address these challenges, we use Qwen3.8-Omni-Flash to explore how to connect source understanding, task planning, tool execution, and result delivery into a complete pipeline. It supports end-to-end, long-horizon workflows such as video editing, translation, film commentary, and content creation, advancing Omni from audio-visual understanding toward autonomous action and task completion.
To this end, we have further expanded Qwen-MM-Plugins with on-demand perception, tool use, and workflow execution for long-form audio and video. We have also open-sourced Qwen-Live Harness as a native runtime for continuous, real-time omnimodal interaction. Together, they address long-horizon workflows and real-time interaction while continuing to expand the capabilities of omnimodal agents alongside the model.
PLUGIN Qwen-MM-Plugins The gateway to audio-visual productivity—connecting multimodal understanding, content creation, and agent harnesses. Get started → HARNESS Qwen-Live Harness The gateway to real-time interaction—connecting audio-visual conversations, task delegation, memory, and context management. Get started →
Long-Form Audio-Visual Understanding#
Qwen3.8-Omni-Flash brings a major upgrade to long-form audio-visual understanding—from controllable descriptions and agentic evidence gathering, to understanding meetings and advancing follow-up tasks, and finally to producing video-centered deep research reports. It does not merely process longer content, but finds relevant evidence more precisely, reasons more deeply, and acts more efficiently.
Controllable Audio-Visual Captioning#
There is no single answer to how a video should be described. Content creation prioritizes narrative, footage retrieval focuses on specific segments, and asset management depends on structure. Different applications need different video descriptions. In Qwen3.8-Omni-Flash, we have upgraded video captioning from answering "what the model saw" to understanding "what the user wants to know." Users can freely specify the subject, time range, level of detail, and output format. For the same video, the model can provide an overview, locate key segments, or analyze character actions, camera shots, lighting, and sound in depth, producing structured results as needed. You define what to look at, how closely to look, and how to present it.
Agentic Long-Form Audio-Visual Understanding#
For videos lasting several hours, conventional approaches require the model to process the entire recording from beginning to end, even when the answer appears in only a few minutes of footage. The native Qwen3.8-Omni-Flash agent starts from the question, independently decides what to watch and listen to, and locates key information through multiple rounds of coarse-to-fine evidence gathering. Without processing every frame, it can focus limited compute and token budgets on the relevant segments, enabling more efficient long-form video understanding. On OmniVideoBench, Agentic Understanding improves accuracy from 63.4 to 67.8 while reducing token consumption from 145,736 to 79,117, a reduction of approximately 45.7%. The table below compares the accuracy and token consumption of Static Understanding and Agentic Understanding on OmniVideoBench:
| Static Understanding | Agentic Understanding | |
|---|---|---|
| Accuracy (↑) | 63.4 | 67.8 |
| Tokens per query (↓) | 145,736 | 79,117 |
Long Meetings: From Minutes to Action#
Multi-participant meetings are among the most complex audio-visual understanding scenarios: speakers take turns and overlap, while identities, references, and discussion topics continuously change. Qwen3.8-Omni-Flash jointly recognizes speakers across audio and video and natively supports up to one hour of audio-visual input. It can perform speaker segmentation, content transcription, and identity alignment end to end. Given a complete meeting video and a request, the model can map participant relationships, generate meeting minutes, identify action items, and analyze project risks, using visual information to resolve references and entity ambiguity in the audio. Combined with agents and tool use, it can also send emails, organize tasks, and even begin coding in response to meeting requirements—moving from understanding a meeting to acting on it.
Conducting Deep Research with Audio and Video#
When users watch a video with a specific question in mind, the answer often extends beyond the video itself. Qwen3.8-Omni-Flash combines the user’s needs with the video content to identify questions worth deeper investigation, organize the key material, and search multimodal sources across the web—including images, videos, and documents. It then produces a video-centered, richly illustrated research report that helps users understand the content and solve practical problems. For example, when a user encounters color fringing around a Photoshop hair cutout, the model can break down the tutorial steps, study the principles behind Multiply and Screen blend modes, compare alternative edge-repair techniques, and explain which approach best fits the user’s situation.
Audio-Visual Production and Editing#
Qwen3.8-Omni-Flash is taking audio-visual agents into a new stage: from understanding sounds and images to independently planning, calling tools, and delivering finished videos, bringing omnimodal intelligence into professional audio-visual content production workflows.
Music2MV#
For music video (MV) creation, Qwen3.8-Omni-Flash can understand the structure, rhythm, mood, vocals, and instrumental changes of a user-provided song in fine detail, informing the design of characters, scenes, and shots. It can also output line-level lyrics with timestamps to align singing, subtitles, and visuals. Combined with creative tools such as Qwen-MM-Plugins, the model supports the complete workflow from music understanding and creative planning to final quality review, demonstrating strong audio-visual understanding, reasoning, and creation capabilities.
Workflow
Next
Final Result 1
Next
Final Result 2
Next
Final Result 3
Next
Short Drama Translation#
Traditional video translation often requires repeatedly switching between transcription, translation, dubbing, and editing platforms. This complicates API calls and workflow coordination and makes it difficult to maintain consistency across character voices, dialogue duration, and visual pacing. With an agent built on Qwen3.8-Omni-Flash, users can describe their needs in a single sentence to perform speaker-aware dialogue recognition, conversational translation, character voice cloning and dubbing, audio remixing, and final quality review. These otherwise fragmented localization steps become a complete workflow, enabling automated delivery of short dramas for international audiences.
Workflow
Next
Final Result 1
Next
Source
Translated Result
Final Result 2
Next
Source
Translated Result
Final Result 3
Next
Source
Translated Result
Long-Form Film Commentary#
Producing commentary videos for full-length films of two or three hours often requires repeatedly watching the film and reconstructing its plot, followed by shot selection, scriptwriting, voiceover, music, and editing. This is a complex and time-consuming process. With an agent built on Qwen3.8-Omni-Flash, users need only provide a film and describe their creative requirements in one sentence to perform long-form video understanding, key-plot extraction, commentary planning, voiceover and music production, editing, rendering, and final quality review. The agent can also intelligently interleave original dialogue with commentary, automatically adjusting speech rate and volume so that narration, original audio, background music, and visuals flow naturally together, creating a more authentic, immersive, and cinematic commentary video.
Workflow
Next
Final Result 1
Next
Final Result 2
Next
From Using Models to Optimizing Models#
Real-world multimodal applications are often complex and cost-sensitive, requiring models to balance quality, latency, compute, and deployment costs.
Customizing smaller models for specific scenarios is therefore an important path to deploying applications at scale. Yet traditional workflows involve data construction, problem diagnosis, multiple rounds of training, and evaluation, making them time-consuming and heavily dependent on human expertise. This time, we extend Qwen3.8-Omni-Flash into model development itself, exploring a new approach in which large models drive research and development while smaller models serve business needs.
We gave Qwen3.8-Omni-Flash a task: improve Qwen2.5-Omni-3B's Sichuan dialect speech recognition within 12 hours and deliver a usable model. It independently selected the WenetSpeech-Chuan evaluation set, fixed the evaluation criteria, and established a baseline. It then listened directly to audio samples, diagnosed problems using the recognition results, and constructed targeted training data. Across four consecutive rounds of experiments, the agent created 3,413 training examples, adjusted its approach based on evaluation feedback, retained effective improvements, and rolled back unsuccessful attempts. Qwen2.5-Omni-3B's character error rate on the same evaluation set ultimately fell from 25.79% to 15.30%, a relative reduction of approximately 40.7%.
This experiment demonstrates another possibility for model evolution: general-purpose multimodal models understand data, plan experiments, and drive iteration, while smaller models acquire specialized capabilities for specific applications. Agents can go beyond using models to help solve practical business problems.
Audio-Visual Information Compression#
Audio and video carry rich information, but their linear, unstructured form makes retrieval and reuse difficult. Qwen3.8-Omni-Flash understands content across sound, visuals, and timelines, using an agentic workflow to extract information, reorganize its structure, and verify results. It transforms the core knowledge and practical experience in long videos into denser information assets that are easier to consume and reuse.
Video2Note#
To turn video knowledge into structured resources, we have open-sourced Video2Note in Qwen-MM-Plugins. Drawing on Qwen3.8-Omni-Flash's joint understanding of speech, visuals, and procedures, it automatically organizes knowledge, breaks down key steps, selects representative frames, and generates PDF notes with corresponding text and images. Automated review and iterative correction further condense hours of video into clear, readable documents that are easy to revisit.
Workflow
Next
Final Result 1
Next
Example Video
Final PDF
Final Result 2
Next
Example Video
Final PDF
Final Result 3
Next
Example Video
Final PDF
Final Result 4
Next
Example Video
Final PDF
Omni Skill Creator#
Videos record not only "how to do something," but also the practical expertise accumulated by specialists. With this in mind, we introduce Omni Skill Creator as a new open-source capability in Qwen-MM-Plugins. It can extract standard operating procedures (SOPs) from demonstrations to perform reusable automated work, or learn tool usage, decision criteria, and key insights from expert instruction. A single demonstration becomes an agent skill that has been verified and evaluated, enabling reusable, shareable skills built from omnimodal content.
Real-Time Audio-Visual Interaction#
Qwen3.8-Omni-Flash is designed for deep understanding and creation with complete audio-visual content. For continuous, low-latency interaction, we further introduce Qwen3.8-Omni-Flash-Realtime. It perceives and responds while receiving live audio-visual streams, and uses real-time context to call tools and execute tasks, taking omnimodal capabilities from “understanding a piece of content” to “participating in an interaction.”
Real-Time Speaking Practice#
Spoken language has no standard input. Accents, vowel and consonant substitutions, and tonal deviations can cause word-for-word transcription to diverge from the intended meaning. Qwen3.8-Omni-Flash-Realtime jointly models pronunciation and semantics, understands nonstandard expressions affected by accents, aligns them with the correct words, and generates standard-pronunciation demonstrations in real time. Across multiple practice turns, the model updates its judgment with new audio, correcting errors that affect understanding while preserving natural rhythm, tone, and emotion.
Omnimodal Spatial Audio Perception#
In real-world spaces, sound provides another coordinate axis beyond vision. Qwen3.8-Omni-Flash-Realtime combines spatial sound with visual information to continuously determine the direction and distance of sound sources while perceiving obstacles, navigable areas, and changes in the scene, making it the first omnimodal model capable of locating targets by sound.
For instructions such as “come over here” or “go see what is making that sound,” the model can isolate voices and target sounds from environmental noise, ground their meaning in the surrounding space, and call tools to perform localization, search, path planning, and navigation—from hearing a target to reaching it.
External Knowledge for Audio-Visual Interaction#
Real-time interaction requires not only low latency, but also the ability to load knowledge and behavior dynamically for each application. Qwen3.8-Omni-Flash-Realtime supports injecting identity settings, expression styles, business knowledge, and interaction rules through Skills, while tool use extends these capabilities into task execution.
In scenarios such as customer service, the model can load brand language and service procedures in real time, understand the user's speech, visuals, and context, generate responses that follow business requirements, and execute actions. The same real-time model can therefore take on different knowledge, roles, and ways of acting.
Benchmark Results#
Omni#
| Qwen3.8-Omni-Flash | Qwen3.5-Omni-Plus | Gemini 3.8 Flash | Seed 2.0 Lite | Muse Spark 1.2 | |
|---|---|---|---|---|---|
| Agentic Omni Intelligence | |||||
WildClawBench-MM Multimodal tool use | 71.0 | 34.5 | 58.9 | 41.9 | -- |
UniClawBench Multimodal tool use | 69.6 | 67.1 | 69.0 | 61.2 | -- |
AgenticVBench Multimodal tool use | 36.8 | 14.5 | 45.0 | 10.0 | -- |
OmniGAIA Web Search | 74.0 | 57.2 | 78.6 | 64.4 | -- |
| General Audio-Visual Capabilities | |||||
DailyOmni Audio-Visual Understanding | 85.1 | 85.1 | 84.0 | 81.4 | 79.6 |
WorldSense Audio-Visual Understanding | 68.5 | 63.9 | 69.6 | 67.3 | 65.0 |
AVUT Audio-Visual Understanding | 86.6 | 85.9 | 88.0 | 81.5 | 82.4 |
JoinAVBench Audio-Visual Understanding | 75.9 | 74.1 | 70.4 | 70.6 | 71.8 |
OmniVideoBench Audio-Visual Reasoning | 63.4 | 53.8 | 65.2 | 58.5 | 62.2 |
Video-MME-v2 Audio-Visual Reasoning | 65.0 | 47.9 | 71.0 | 64.9 | -- |
LVOmniBench Long Video Reasoning | 63.3 | 53.2 | 70.7 | -- | -- |
OmniCloze Audio-Visual Caption | 63.2 | 64.2 | 60.9 | 56.3 | 65.3 |
OmniCap-IF Audio-Visual Caption | CSR 80.6 ISR 28.2 | CSR 72.1 ISR 14.1 | CSR 81.9 ISR 28.3 | CSR 74.6 ISR 18.1 | CSR 77.9 ISR 26.8 |
QIVD Audio-Visual Interaction | 69.6 | 65.6 | 69.1 | 62.0 | 62.0 |
StreamingBench Audio-Visual Interaction | 80.8 | 57.1 | 79.9 | 77.2 | 77.8 |
| General Audio Capabilities | |||||
AliMeeting Test Multi-Speaker ASR (DER | cpWER, ↓) | 3.4 | 17.2 | 88.1 | 89.6 | 72.6 | 53.1 | 75.1 | 76.1 | 93.7 | 92.7 |
AISHELL-4 Multi-Speaker ASR (DER | cpWER, ↓) | 2.8 | 11.2 | 100.0 | 100.0 | 66.4 | 56.9 | 64.8 | 64.2 | 91.3 | 86.0 |
MagicData-RAMC Multi-Speaker ASR (DER | cpWER, ↓) | 5.7 | 14.1 | 98.4 | 97.1 | 67.9 | 33.8 | 43.4 | 35.1 | 82.1 | 75.3 |
MLC-SLM (en) Multi-Speaker ASR (DER | cpWER, ↓) | 4.0 | 14.2 | 68.6 | 63.9 | 60.8 | 26.6 | 40.4 | 45.5 | 74.3 | 52.9 |
WenetSpeech (Net) ASR (WER, ↓) | 4.8 | 3.7 | 14.2 | 4.3 | 68.2 |
WenetSpeech (Meeting) ASR (WER, ↓) | 4.6 | 4.8 | 16.7 | 4.7 | 42.6 |
FLEURS-ASR Multilingual ASR (WER, ↓) | 9.3 | 7.2 | 7.9 | 32.1 | 23.6 |
FLEURS-S2TT Multilingual S2TT (BLEU) | 31.8 | 32.2 | 33.0 | 24.8 | 28.8 |
SpotSoundBench Audio Grounding | 67.2 | 64.2 | 39.7 | 59.6 | 16.9 |
MMAU Audio Understanding | 81.8 | 81.9 | 76.9 | 77.2 | 63.5 |
MMAR Audio Understanding | 79.8 | 79.8 | 78.5 | 77.7 | 67.3 |
MMSU Audio Understanding | 82.1 | 83.0 | 83.3 | 80.2 | 59.9 |
LongAudioSpan Long Audio Reasoning | Accuracy 82.7 Rubric 71.8 Chain 48.2 | Accuracy 74.4 Rubric 49.8 Chain 45.1 | Accuracy 79.3 Rubric 65.5 Chain 64.6 | -- | -- |
MuchoMusic-RUL Music Understanding | 72.6 | 71.6 | 53.7 | 61.7 | 40.1 |
HumMusQA Music Understanding | 75.8 | 75.5 | 71.2 | 66.0 | 63.3 |
MusTBench Music Understanding | 50.6 | 49.1 | 40.3 | 44.0 | 29.4 |
Audio MultiChallenge Audio Interaction | 71.5 | 57.6 | 71.9 | 63.4 | 57.9 |
WildSpeech Audio Interaction | 74.3 | 75.7 | 76.4 | 74.5 | 73.4 |
VoiceBench Audio Interaction | 91.6 | 92.9 | 92.3 | 84.1 | 79.8 |
1. Evaluation harnesses for Agentic Omni Intelligence: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, and OmniGAIA uses no harness. For WildClawBench-MM, we evaluate only the multimodal tasks in WildClawBench that involve images, video, or audio.
2. FLEURS: ASR and S2TT evaluation results both cover the following 60 languages: Chinese (Mandarin), English, Cantonese, Arabic, German, French, Spanish, Portuguese, Indonesian, Italian, Korean, Russian, Thai, Vietnamese, Japanese, Turkish, Hindi, Malay, Dutch, Urdu, Norwegian, Swedish, Danish, Hebrew, Finnish, Polish, Icelandic, Czech, Filipino, Persian, Greek, Afrikaans, Asturian, Belarusian, Bulgarian, Bengali, Bosnian, Catalan, Cebuano, Estonian, Galician, Gujarati, Croatian, Hungarian, Javanese, Kazakh, Kannada, Kyrgyz, Latvian, Macedonian, Malayalam, Marathi, Punjabi, Romanian, Slovak, Slovenian, Swahili, Tajik, Azerbaijani, and Ukrainian.
3. Empty cells (--): scores are not yet available or are not applicable.
Agentic Omni Understanding#
Key information in long audio and video recordings is often scattered across different segments, while complex questions require multiple steps of reasoning across sound and images. Agentic Omni Understanding enables the model to start from the question, plan its approach, call tools, and progressively locate and verify evidence. By focusing computation on relevant content, it improves the accuracy and efficiency of long-form audio-visual understanding. To evaluate this capability, we compare Qwen3.8-Omni-Flash and Gemini 3.8 Flash on OmniVideoBench, Video-MME-v2, and LVOmniBench under two settings: Static, where the model directly interprets the input, and an agent mode using Qwen Code. This comparison shows how introducing agent workflows affects each model’s performance.
| Qwen3.8-Omni-Flash (Static) | Qwen3.8-Omni-Flash (Qwen Code) | Gemini 3.8 Flash (Static) | Gemini 3.8 Flash (Qwen Code) | |
|---|---|---|---|---|
OmniVideoBench Audio-Visual Reasoning | 63.4 | 67.8 | 65.2 | 70.1 |
Video-MME-v2 Audio-Visual Reasoning | 65.0 | 71.3 | 71.0 | 72.7 |
LVOmniBench Long Video Reasoning | 63.3 | 73.6 | 70.7 | 70.7 |
Text#
| Qwen3.8-Omni-Flash | Qwen3.8- Flash | Qwen3.8- 27B | Qwen3.7- Plus | DeepSeek-V4-Flash-0731 | Claude-Opus-4.6 (Max) | |
|---|---|---|---|---|---|---|
| Coding and Agent | ||||||
DeepSWE 1.1 Long-horizon software engineering | 57.8 | 58.7 | 42.2 | 16.5 | 54.4 | -- |
SWE-bench Pro Long-horizon software engineering | 63.3 | 62.5 | 61.7 | 55.8 | 56.0 | 53.4 |
SWE-bench Multilingual Multilingual software engineering | 80.5 | 81.0 | 73.8 | 75.8 | -- | 77.5 |
NL2Repo-Bench Repo-level code generation | 48.9 | 48.1 | 42.3 | 41.1 | 54.2 | 47.6 |
CoWorkBench Long-horizon office work | 75.3 | 73.9 | 70.7 | 65.1 | 45.1 | 68.2 |
| General Text Capabilities | ||||||
IFBench Instruction following | 81.5 | 81.3 | 79.5 | 79.1 | 79.2 | 62.5 |
GPQA Diamond Scientific reasoning | 91.0 | 91.7 | 89.2 | 90.3 | 90.8 | 91.3 |
HLE Multidisciplinary reasoning | 36.5 | 35.9 | 30.8 | 34.7 | 33.8 | 40.0 |
LiveCodeBench v6 Competitive coding | 92.6 | 91.9 | 90.3 | 89.6 | 90.6 | 88.8 |
1. DeepSWE 1.1: evaluated with the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, 256K context window. We report the highest score across the two harnesses; notably, Qwen3.8-Flash-Next performs best on mini-SWE-agent.
2. SWE-bench Pro: except for Claude-Opus-4.6 (Max), for which we report the officially published score, all models are evaluated with the Claude Code harness, temp=1.0, top_p=0.95, 256K context window. Problematic tasks were corrected and all baseline models were re-evaluated on the refined benchmark.
3. SWE-bench Multilingual: evaluated with the mini-SWE-agent harness, temp=1.0, top_p=0.95, 256K context window.
4. NL2Repo-Bench: evaluated with the Claude Code harness. To prevent reward hacking, we disable Bash commands that attempt to access the specific repository, such as pip download, pip install and git clone.
5. CoWorkBench: an in-house cowork benchmark for evaluating long-horizon office and productivity agent tasks across computer science, finance, law, medical and other productivity domains.
6. HLE: judged by GPT-4o.
7. Empty cells (--): scores are not yet available or are not applicable.
Vision#
| Qwen3.8-Omni-Flash | Qwen3.8- Flash | Qwen3.8- 27B | Qwen3.7- Plus | Claude-Opus-4.6 (Max) | |
|---|---|---|---|---|---|
| Agentic Vision Intelligence | |||||
ClawEval-MM Multimodal tool use | Pass@3 60.4 Average 61.9 | Pass@3 64.4 Average 60.4 | Pass@3 57.4 Average 56.9 | Pass@3 57.4 Average 60.1 | Pass@3 52.5 Average 54.7 |
AndroidWorld Mobile use | 87.1 | 84.5 | 81.9 | 81.0 | 62.0 |
Vision2Web Visual web development | 62.9 | 64.0 | 62.9 | 42.1 | -- |
| General Vision Capabilities | |||||
ERQA Embodied intelligence | 71.0 | 72.3 | 65.5 | 69.8 | 40.8 |
LVBench Long video understanding | 76.9 | 76.6 | 72.4 | 76.2 | 63.0 |
RealWorldQA Real-world perception | 87.7 | 88.5 | 85.9 | 86.9 | 73.9 |
MathVision Visual math problem solving | Without CI 91.8 With CI 96.2 | Without CI 90.6 With CI 95.7 | Without CI 90.0 With CI 94.6 | Without CI 90.3 With CI 88.4 | Without CI 65.5 With CI -- |
CharXiv (RQ) Scientific chart analysis | Without CI 83.5 With CI 91.4 | Without CI 84.6 With CI 90.6 | Without CI 83.7 With CI 90.2 | Without CI 85.8 With CI 85.9 | Without CI 66.0 With CI -- |
1. ClawEval-MM: scores are reported as "pass@3 / average score". Pass@3 measures the percentage passed in at least one of three trials, and the average score is the mean score across the three trials.
2. RecreationBench: an in-house long-horizon application-recreation benchmark for evaluating hybrid-agent abilities spanning five platforms — desktop (Ubuntu, macOS, Windows), mobile (Android) and web.
3. OSWorld 2.0: scores are reported as "binary / partial". The binary score is the percentage of tasks that receive the full task reward, while the partial score aggregates the partial rewards obtained across all tasks.
4. Vision2Web: scores are reported as the average over the frontend, webpage and website categories, using the Claude Code harness and judged by gpt-5.4-2026-03-05.
5. MathVision, CharXiv (RQ): scores are reported as "without CI / with CI". A small number of incorrect ground-truth annotations in MathVision were corrected after manual verification. Our model's score is evaluated using a fixed prompt, e.g. "Please reason step by step, and put your final answer within \boxed{}." For other models, we report the higher score between runs with and without the \boxed{} formatting.
6. Empty cells (--) indicate scores not yet available or not applicable.
Throughput & Latency#
The following results show the observed throughput and latency of the Qwen3.8-Omni-Flash-Realtime API under different input conditions, reflecting the performance users experience in production environments.
| Input Scenario | Text Output TPS (Tokens/s) | Time to First Token (ms) | Time to First Audio Packet (ms) | Audio Generation RTF |
|---|---|---|---|---|
| Realtime API Performance | ||||
| Audio 6s | 84.87 | 591.26 | 978.36 | 0.1538 |
| Audio 12s | 83.18 | 604.80 | 982.74 | 0.1537 |
| Audio 20s | 81.06 | 617.98 | 1026.39 | 0.1538 |
| Audio-Visual 6s | 84.89 | 837.96 | 1214.73 | 0.1524 |
| Audio-Visual 12s | 84.33 | 911.85 | 1268.07 | 0.1527 |
| Audio-Visual 20s | 83.00 | 981.01 | 1350.49 | 0.1528 |
Supported Languages#
| Capability | Languages | Chinese Dialects |
|---|---|---|
| Speech Recognition | 74 languages: Afrikaans, Arabic, Asturian, Azerbaijani, Basque, Belarusian, Bengali, Bosnian, Bulgarian, Cantonese, Catalan, Cebuano, Chinese, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Filipino, Finnish, French, Galician, Georgian, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Interlingua, Italian, Japanese, Javanese, Kannada, Kazakh, Korean, Kyrgyz, Lingala, Latvian, Lithuanian, Macedonian, Malay, Malayalam, Maltese, Maori, Marathi, Mongolian, Norwegian Bokmål, Norwegian Nynorsk, Oriya, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tajiki, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Uyghur, and Vietnamese | 39 dialects: Northeastern Mandarin, Guizhou dialect, Guangdong Cantonese, Henan dialect, Hong Kong Cantonese, Shanghainese, Shaanxi dialect, Tianjin dialect, Taiwanese Mandarin, Yunnan dialect, Anhui dialect, Fujian dialect, Gansu dialect, Guangdong Mandarin, Hubei dialect, Hunan dialect, Jiangxi dialect, Shandong dialect, Shanxi dialect, Sichuanese, Guangxi dialect, Hainan dialect, Chongqing dialect, Changsha dialect, Hangzhou dialect, Hefei dialect, Yinchuan dialect, Zhengzhou dialect, Shenyang dialect, Wenzhou dialect, Wuhan dialect, Kunming dialect, Taiyuan dialect, Nanchang dialect, Jinan dialect, Lanzhou dialect, Nanjing dialect, Hakka, and Southern Min |
| Speech Generation | 29 languages: Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian, Thai, Indonesian, Arabic, Vietnamese, Turkish, Finnish, Polish, Hindi, Dutch, Czech, Urdu, Tagalog, Swedish, Danish, Hebrew, Icelandic, Malay, Norwegian, and Persian | 7 dialects: Sichuanese, Beijing dialect, Tianjin dialect, Nanjing dialect, Shaanxi dialect, Cantonese, and Southern Min |
Getting Started with Qwen3.8-Omni-Flash#
API Usage#
Qwen3.8-Omni-Flash officially supports reasoning_effort to adjust reasoning depth and control costs:
xhigh(default): for complex tasks that require in-depth analysis.medium: balances accuracy and speed.low: efficient reasoning optimized for speed and cost.
In addition, preserve_thinking is enabled by default across all scenarios for the best out-of-the-box experience.
Qwen3.8-Omni-Flash supports industry-standard protocols, including OpenAI-compatible Chat Completions and Responses APIs. Examples follow:
"""
Environment variables:
DASHSCOPE_API_KEY: Your API Key from https://platform.qianwenai.com/home
DASHSCOPE_BASE_URL: (optional) Base URL for compatible-mode API.
- Beijing: https://dashscope.aliyuncs.com/compatible-mode/v1
- Singapore: https://dashscope-intl.aliyuncs.com/compatible-mode/v1
"""
from openai import OpenAI
import os
api_key = os.environ.get("DASHSCOPE_API_KEY")
if not api_key:
raise ValueError(
"DASHSCOPE_API_KEY is required. "
"Set it via: export DASHSCOPE_API_KEY='your-api-key'"
)
client = OpenAI(
api_key=api_key,
base_url=os.environ.get(
"DASHSCOPE_BASE_URL",
"https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
),
)
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241022/emyrja/dog_and_girl.jpeg"
},
},
{
"type": "input_audio",
"input_audio": {
"data": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250211/tixcef/cherry.wav",
"format": "wav"
},
},
{"type": "text", "text": "Please describe the image and tell me what is being said in the audio."},
],
},
]
completion = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=messages,
extra_body={
"enable_thinking": True,
# "preserve_thinking": True,
},
reasoning_effort="xhigh", # supported levels are xhigh, medium, and low
stream=True,
)
reasoning_content = ""
answer_content = ""
is_answering = False
print("\n" + "=" * 20 + "Reasoning" + "=" * 20 + "\n")
for chunk in completion:
if not chunk.choices:
print("\nUsage:")
print(chunk.usage)
continue
delta = chunk.choices[0].delta
if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None:
if not is_answering:
print(delta.reasoning_content, end="", flush=True)
reasoning_content += delta.reasoning_content
if hasattr(delta, "content") and delta.content:
if not is_answering:
print("\n" + "=" * 20 + "Answer" + "=" * 20 + "\n")
is_answering = True
print(delta.content, end="", flush=True)
answer_content += delta.content
Best Practices for Controllable Audio-Visual Captioning#
Content organization and level of detail: Organize video descriptions, OCR extraction, audio descriptions, and speech transcripts into sections covering visuals, speech, music, sound effects, and ambient sounds. Preserve the original text, speaker identities, and corresponding time ranges to make retrieval and verification easier. Specify a target length in the prompt to control detail, for example, Describe the video in approximately 2000–3000 words., and adjust it to the video’s duration and information density.
Prompt Example
Provide a detailed description of the video.
Make sure your description covers every one of the following dimensions:
Visual
- Subjects and characters: appearance, clothing, gender/age cues, identity, distinctive features
- Actions and events in chronological order, and how the scene evolves over time
- Setting and background: location, environment, time of day
- Spatial layout and relations between subjects/objects; counts and quantities
- On-screen text: captions, titles, subtitles, logos, UI — exact content and appearance
- Visual style: colors, lighting, camera shots, angles, and camera movement
Audio
- Speech: the exact spoken content, transcribed verbatim
- Speakers: who is speaking (mapped to the on-screen person or voice-over), with accent, tone, gender/age cues
- Speaking state: prosody, emotion, volume, and speaking style
- Music: presence, genre/mood, and lyrics if any
- Sound effects and ambient/background sounds
- Non-speech vocalizations: laughter, crying, applause, etc.
Audio-visual correspondence
- Which speech or sound aligns with which on-screen person or visual event
- The timing of each event, expressed with timestamps
It should explicitly include three sections:
1. A structured chronological storyline of **every noticeable audio and visual details**
2. A structured list of all visible text. For each text element, include start timestamp, end timestamp, the exact text content, the appearance characteristics. If no text appears, explicitly state so.
3. A structured speech-to-text transcription, include speaker (corresponding to the character or voice-over in Section 1, including their accent and tone), exact spoken content, start timestamp, end timestamp, and speaking state (prosody, emotion, and style). If no speech appears, explicitly state so.
Aside from these three required sections, you are free to organize any additional content in any way you find helpful. This additional content can include global information about the entire video or localized information about specific moments. You may choose the topic of this extra content freely.
Rules:
- Add as much descriptive detail as possible.
- Do not use Markdown bold formatting.
- Carefully look at frames and listen to the audio, making sure no detail is overlooked.
Output Format:
```
## Storyline
<xx:xx.xxx> - <xx:xx.xxx>
<an unstructured long paragraph in natural language describing what happened during this period, blending both audio and video details.>
<xx:xx.xxx> - <xx:xx.xxx>
<an unstructured long paragraph in natural language describing what happened during this period, blending both audio and video details.>
<xx:xx.xxx> - <xx:xx.xxx>
<an unstructured long paragraph in natural language describing what happened during this period, blending both audio and video details.>
...
## Visible Text
<xx:xx.xxx> - <xx:xx.xxx>
“<element>”: <appearance>
“<element>”: <appearance>
<xx:xx.xxx> - <xx:xx.xxx>
“<element>”: <appearance>
“<element>”: <appearance>
“<element>”: <appearance>
<xx:xx.xxx> - <xx:xx.xxx>
“<element>”: <appearance>
...
## Speakers and Transcript
Speaker profiles:
<speaker> - <profile>
<speaker> - <profile>
<speaker> - <profile>
...
<xx:xx.xxx> - <xx:xx.xxx>
Speaker: <speaker>
State: <description>
Content: “<content>”
<xx:xx.xxx> - <xx:xx.xxx>
Speaker: <speaker>
State: <description>
Content: “<content>”
<xx:xx.xxx> - <xx:xx.xxx>
Speaker: <speaker>
State: <description>
Content: “<content>”
...
## <another section>
<paragraphs>
## <another section>
<paragraphs>
...
```
Structured output and schema constraints: Include task instructions and a complete JSON Schema in the prompt, specifying field meanings, types, required fields, and constraints to support programmatic parsing and downstream use.
Prompt Example
Describe the audio and visual content in detail in English, organized into scenes and events, following the JSON Schema below.
All timestamps must be relative to the beginning of the video. End times must not precede start times or exceed the video duration. Each event must fall within the time range of its parent scene.
Include only information directly supported by the audio or video. Do not guess or invent details. Do not infer causality merely because a sound and an action occur at the same time.
Return only valid JSON, without Markdown fences or commentary.
JSON Schema:
{
"$defs": {
"Event": {
"additionalProperties": false,
"properties": {
"time_range": {
"$ref": "#/$defs/TimeRange",
"description": "Time range of the event"
},
"participants": {
"description": "People, animals, or objects involved, named by observable features; use consistent names for the same participant",
"items": {
"type": "string"
},
"title": "Participants",
"type": "array"
},
"action": {
"description": "Specific actions, interactions, and observable outcomes",
"title": "Action",
"type": "string"
},
"sounds": {
"description": "Sounds heard during the event; use an empty list if none are discernible",
"items": {
"type": "string"
},
"title": "Sounds",
"type": "array"
}
},
"required": [
"time_range",
"participants",
"action",
"sounds"
],
"title": "Event",
"type": "object"
},
"Scene": {
"additionalProperties": false,
"properties": {
"time_range": {
"$ref": "#/$defs/TimeRange",
"description": "Time range of the scene"
},
"setting": {
"description": "Environment, spatial layout, and main visual features",
"title": "Setting",
"type": "string"
},
"events": {
"description": "Events in chronological order; use an empty list if there are none",
"items": {
"$ref": "#/$defs/Event"
},
"title": "Events",
"type": "array"
}
},
"required": [
"time_range",
"setting",
"events"
],
"title": "Scene",
"type": "object"
},
"TimeRange": {
"additionalProperties": false,
"properties": {
"start_seconds": {
"description": "Start time in seconds relative to the beginning of the video",
"minimum": 0,
"title": "Start Seconds",
"type": "number"
},
"end_seconds": {
"description": "End time in seconds; must not precede the start time",
"minimum": 0,
"title": "End Seconds",
"type": "number"
}
},
"required": [
"start_seconds",
"end_seconds"
],
"title": "TimeRange",
"type": "object"
}
},
"additionalProperties": false,
"properties": {
"summary": {
"description": "An overview of the main content of the video",
"title": "Summary",
"type": "string"
},
"scenes": {
"description": "Scenes in chronological order; group continuous footage with a consistent setting into one scene",
"items": {
"$ref": "#/$defs/Scene"
},
"title": "Scenes",
"type": "array"
}
},
"required": [
"summary",
"scenes"
],
"title": "CaptionResult",
"type": "object"
}
Installing Qwen-MM-Plugins in an Agent Harness#
Qwen-MM-Plugins is a multimodal plugin suite for agent harnesses. It gives agents the ability to understand images, audio, video, and documents, as well as maintain memory for long-form video and create content. It supports agent harnesses including Codex, Claude Code, Qwen Code, Gemini CLI, Qoder, CodeBuddy, and OpenClaw. We also welcome contributions from community developers.
You can enter the following request directly in your usual office agent:
Help me install the core, api, and omni-related plugins from https://github.com/QwenLM/Qwen-MM-Plugins.

Alternatively, install from the command line:
# Run the official guided installer:
curl -fsSL https://raw.githubusercontent.com/QwenLM/Qwen-MM-Plugins/main/install.sh | bash
In the menu, select:
- Install
- The agent harness you use
- The Omni plugins you need
| Plugin | Capability |
|---|---|
core | Read images, video frames, PDFs, Office documents, code, data, and 3D files |
api | Image understanding, OCR, object localization, audio-visual transcription, speaker diarization, event analysis, and image segmentation |
omni-chatcut | Create MVs, long-form film commentary, and video speech translations |
omni-video2note | Turn video tutorials into PDF notes with key screenshots |
omni-skill-creator | Turn instructional videos, screen recordings, or operation demonstrations into reusable Agent Skill.md files |
omni-memory | Build memory for people, dialogue, sounds, and events in long-form videos |
After installation, restart the agent harness or create a new task.
@song.mp3 Generate a complete MV based on the song's rhythm and content.
@short_drama.mp4 Translate the video into English while preserving the original speakers' voice characteristics where possible.
@movie.mp4 Create a film commentary video with Chinese narration and subtitles.
@weekly_report_sop.mp4 Turn this screen recording of writing a weekly report into a Skill.md file.
@tutorial.mp4 Turn the tutorial into PDF notes with key screenshots, timestamps, and step-by-step instructions.
@documentary.mp4 Build audio-visual memory that records people, dialogue, sounds, and important events.
Getting Started with Qwen3.8-Omni-Flash-Realtime#
API Usage#
Qwen3.8-Omni-Flash-Realtime supports connections over WebSocket and WebRTC. Running the basic examples below opens the camera and microphone for a real-time audio-visual conversation. We recommend using headphones.
# Run pip install websocket-client pyaudio dashscope opencv-python -U to install dependencies
"""
Environment variables:
DASHSCOPE_API_KEY: Your API Key from https://platform.qianwenai.com/home
DASHSCOPE_BASE_URL: (optional) Base URL for realtime API.
- Beijing: wss://dashscope.aliyuncs.com/api-ws/v1/realtime
- Singapore: wss://dashscope-intl.aliyuncs.com/api-ws/v1/realtime
"""
import os
import base64
import time
import pyaudio
import cv2
from dashscope.audio.qwen_omni import MultiModality, AudioFormat,OmniRealtimeCallback,OmniRealtimeConversation
import dashscope
url = f'wss://dashscope-intl.aliyuncs.com/api-ws/v1/realtime'
dashscope.api_key = os.getenv('DASHSCOPE_API_KEY')
# Determine the voice
voice = 'Tina'
# Determine the model
model = 'qwen3.8-omni-flash-realtime'
# Determine the model role
instructions = "You are Qwen-Omni, a helpful assistant."
video_fps = 1
video_size = (1280, 720)
class SimpleCallback(OmniRealtimeCallback):
def __init__(self, pya):
self.pya = pya
self.out = None
def on_open(self):
# Initialize audio output stream
self.out = self.pya.open(
format=pyaudio.paInt16,
channels=1,
rate=24000,
output=True
)
def on_event(self, response):
if response['type'] == 'response.audio.delta':
# Play audio
self.out.write(base64.b64decode(response['delta']))
elif response['type'] == 'conversation.item.input_audio_transcription.delta':
# Streaming preview: text is the confirmed prefix, stash is the confirmed suffix
preview = response.get('text', '') + response.get('stash', '')
print(f"\r[User] {preview}", end='', flush=True)
elif response['type'] == 'conversation.item.input_audio_transcription.completed':
# Transcription completed, print the final text and a new line
print(f"\r[User] {response['transcript']}")
elif response['type'] == 'response.audio_transcript.done':
# Print the assistant's response text
print(f"[LLM] {response['transcript']}")
# 1. Initialize audio device
pya = pyaudio.PyAudio()
# 2. Create callback function and session
callback = SimpleCallback(pya)
conv = OmniRealtimeConversation(model=model, callback=callback, url=url)
# 3. Establish connection and configure session
conv.connect()
conv.update_session(output_modalities=[MultiModality.AUDIO, MultiModality.TEXT], voice=voice, instructions=instructions)
# 4. Initialize audio input stream
mic = pya.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True)
camera = cv2.VideoCapture(0)
next_frame_at = 0.0
# 5. Main loop to process audio and video input
try:
if not camera.isOpened():
raise RuntimeError("Cannot open camera 0.")
camera.set(cv2.CAP_PROP_FRAME_WIDTH, video_size[0])
camera.set(cv2.CAP_PROP_FRAME_HEIGHT, video_size[1])
camera.set(cv2., 1)
print(f"Conversation started with camera ({video_fps} fps, {video_size[0]}x{video_size[1]}), speak into the microphone (Ctrl+C to exit)...")
while True:
audio_data = mic.read(3200, exception_on_overflow=False)
conv.append_audio(base64.b64encode(audio_data).decode())
success, frame = camera.read()
if not success:
raise RuntimeError("Cannot read a camera frame.")
if time.monotonic() >= next_frame_at:
frame = cv2.resize(frame, video_size)
success, image = cv2.imencode('.jpg', frame, [cv2.IMWRITE_JPEG_QUALITY, 80])
if not success:
raise RuntimeError("Cannot encode a camera frame.")
conv.append_video(base64.b64encode(image).decode())
next_frame_at = time.monotonic() + 1 / video_fps
time.sleep(0.01)
except KeyboardInterrupt:
pass
finally:
# Clean up resources
camera.release()
conv.close()
mic.close()
if callback.out:
callback.out.close()
pya.terminate()
print("\nConversation ended")
Qwen-Live Harness#
Qwen-Live Harness is a comprehensive open-source harness designed around the Qwen3.8-Omni-Flash-Realtime API. It can be installed with a single command and integrated into mainstream agent workflows. It supports task delegation, proactive interaction, long-term memory, and context management, and welcomes contributions from the community.

Figure 2. Qwen-Live Harness Interaction Framework.
Install and get started:
npm install -g qwen-live-harness
qwen-live-harness init
qwen-live-harness
Citation#
@misc{qwen38omniflash,
title = {Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.},
url = {https://qwen.ai/blog?id=qwen3.8-omni-flash},
author = {{Qwen Team}},
month = {September},
year = {2026}
}
来源:Qwen:Blog Retrieval(API) · qwen.ai