跳到正文
北京时间
原文
Qwen:Blog Retrieval(API)· QwenTeam·· 12 天前精选AI 评分78

Qwen 发布原生全模态模型 Qwen3.8-Omni-Flash,主打音视频智能体任务交付

Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.

AI 导读

Qwen 发布下一代原生全模态模型 Qwen3.8-Omni-Flash,支持文本、图像、音频和视频输入及 1M token 上下文窗口,29 项评测平均分较 Qwen3.5-Omni-Plus 提升超过 25%,音频输入每小时价格下降超过 98%,音视频输入每小时价格下降超过 93%。

推荐理由

官方发布同时开放 API 和开源插件与运行时,读者可以据此评估音视频智能体在剪辑、会议、实时交互等工作流中的落地方式。

正文 · 原文
Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.

QWEN-LIVE HARNESS

QWEN-MM-PLUGINS

QWEN3.8-OMNI-FLASH API

QWEN3.8-OMNI-FLASH-REALTIME API

Introduction#

Today, we are launching Qwen3.8-Omni-Flash, our next-generation native omnimodal model. Its core objective is to strengthen agent capabilities in real-world productivity scenarios, advancing omnimodal models from “understanding omnimodal content” to “planning tasks, calling tools, and completing creative work.” Building on general agentic capabilities in coding, text-based knowledge work, and GUI operation, Qwen3.8-Omni-Flash further extends agentic applications centered on audio and video, delivering strong results across workflows such as video editing, music video creation, film production and commentary, audio-visual summarization, and real-time conversations.

Qwen3.8-Omni-Flash and its applications in production

Figure 1. Qwen3.8-Omni-Flash and its applications in production.

  • Qwen3.8-Omni-Flash — now available on the Qianwen AI Platform:
    • Text, image, audio, and video inputs with a 1M-token context window.

Qwen3.8-Omni-Flash supports a 1M-token context window while maintaining text performance comparable to a text-only model of the same size and delivering significant improvements in omnimodal capabilities. Across 29 evaluations1, its average score improves by more than 25% over Qwen3.5-Omni-Plus; the API price per hour of audio input decreases by more than 98%, and the price per hour of audio-visual input decreases by more than 93%2. For audio-visual agents, coding, and long-horizon tasks, the model improves by 36.5 points on WildClawBench-MM and 22.3 points on AgenticVBench, while scoring a strong 69.6 on UniClawBench. Its core capabilities also improve significantly in long-form audio and audio-visual understanding, audio-visual reasoning, audio-visual captioning, and multi-speaker recognition. For example, it gains 8.3 points on LongAudioSpan and 9.6 points on OmniVideoBench; its OmniCap-IF CSR and ISR improve by 8.5 and 14.1 points, respectively; and its AliMeeting DER and cpWER decrease from 88.11 / 89.61 to 3.35 / 17.18. By scaling data, context, and agentic environments, Qwen3.8-Omni-Flash achieves audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash. These advances also mean that audio and video are evolving from perceptual inputs into core media through which agents understand their environment, reason, and execute tasks.

1. The scope includes audio reasoning benchmarks: AliMeeting-test, AISHELL-4, MagicData-RAMC, MLC-SLM (en), WenetSpeech (Net | Meeting), FLEURS-60 ASR, FLEURS-60 S2TT, SpotSoundBench, MMAU, MMAR, MMSU, MuchoMusic-RUL, HumMusQA, MusTBench, Audio-MultiChallenge, WildSpeech, and VoiceBench; audio-visual reasoning benchmarks: DailyOmni, WorldSense, AVUT, JointAVBench, OmniCloze, OmniCap-IF, QIVD, OmniVideoBench, and StreamingBench; and audio-visual agent benchmarks: WildClawBench-MM, UniClawBench, and OmniGAIA.
2. Pricing methodology: hourly audio or audio-visual input prices are estimated as 30 times the input cost of two minutes of source material; audio-visual input uses 720p at 1 fps. Gemini 3.8 Flash uses media_resolution=high, Seed 2.0 Lite uses max_frame_tokens=384, and all other API parameters use their default values. Text input and output prices are in CNY per 1M tokens; Gemini and Muse prices are converted from USD at an exchange rate of 1 USD = 6.7191 CNY.

Qwen3.8-Omni-Flash pricing and representative benchmark comparison

Audio and video are important media for bringing agents into real-world productivity scenarios, but they also introduce a new set of system-level challenges. Long-form audio and video are costly to store, transmit, and process across multiple rounds of inference; existing agent harness frameworks lack native support for these modalities; and workflows that connect omnimodal understanding with end-to-end task execution are still at an early stage. Addressing these challenges requires models, harness tools, and runtime environments to evolve together.

To address these challenges, we use Qwen3.8-Omni-Flash to explore how to connect source understanding, task planning, tool execution, and result delivery into a complete pipeline. It supports end-to-end, long-horizon workflows such as video editing, translation, film commentary, and content creation, advancing Omni from audio-visual understanding toward autonomous action and task completion.

To this end, we have further expanded Qwen-MM-Plugins with on-demand perception, tool use, and workflow execution for long-form audio and video. We have also open-sourced Qwen-Live Harness as a native runtime for continuous, real-time omnimodal interaction. Together, they address long-horizon workflows and real-time interaction while continuing to expand the capabilities of omnimodal agents alongside the model.

PLUGIN Qwen-MM-Plugins The gateway to audio-visual productivity—connecting multimodal understanding, content creation, and agent harnesses. Get started → HARNESS Qwen-Live Harness The gateway to real-time interaction—connecting audio-visual conversations, task delegation, memory, and context management. Get started →

Long-Form Audio-Visual Understanding#

Qwen3.8-Omni-Flash brings a major upgrade to long-form audio-visual understanding—from controllable descriptions and agentic evidence gathering, to understanding meetings and advancing follow-up tasks, and finally to producing video-centered deep research reports. It does not merely process longer content, but finds relevant evidence more precisely, reasons more deeply, and acts more efficiently.

Controllable Audio-Visual Captioning#

There is no single answer to how a video should be described. Content creation prioritizes narrative, footage retrieval focuses on specific segments, and asset management depends on structure. Different applications need different video descriptions. In Qwen3.8-Omni-Flash, we have upgraded video captioning from answering "what the model saw" to understanding "what the user wants to know." Users can freely specify the subject, time range, level of detail, and output format. For the same video, the model can provide an overview, locate key segments, or analyze character actions, camera shots, lighting, and sound in depth, producing structured results as needed. You define what to look at, how closely to look, and how to present it.

视频 · 前往原文观看

Agentic Long-Form Audio-Visual Understanding#

For videos lasting several hours, conventional approaches require the model to process the entire recording from beginning to end, even when the answer appears in only a few minutes of footage. The native Qwen3.8-Omni-Flash agent starts from the question, independently decides what to watch and listen to, and locates key information through multiple rounds of coarse-to-fine evidence gathering. Without processing every frame, it can focus limited compute and token budgets on the relevant segments, enabling more efficient long-form video understanding. On OmniVideoBench, Agentic Understanding improves accuracy from 63.4 to 67.8 while reducing token consumption from 145,736 to 79,117, a reduction of approximately 45.7%. The table below compares the accuracy and token consumption of Static Understanding and Agentic Understanding on OmniVideoBench:

Note: Agentic mode preserves context across turns.
Static UnderstandingAgentic Understanding
Accuracy (↑)63.467.8
Tokens per query (↓)145,73679,117
视频 · 前往原文观看

Long Meetings: From Minutes to Action#

Multi-participant meetings are among the most complex audio-visual understanding scenarios: speakers take turns and overlap, while identities, references, and discussion topics continuously change. Qwen3.8-Omni-Flash jointly recognizes speakers across audio and video and natively supports up to one hour of audio-visual input. It can perform speaker segmentation, content transcription, and identity alignment end to end. Given a complete meeting video and a request, the model can map participant relationships, generate meeting minutes, identify action items, and analyze project risks, using visual information to resolve references and entity ambiguity in the audio. Combined with agents and tool use, it can also send emails, organize tasks, and even begin coding in response to meeting requirements—moving from understanding a meeting to acting on it.

视频 · 前往原文观看

Conducting Deep Research with Audio and Video#

When users watch a video with a specific question in mind, the answer often extends beyond the video itself. Qwen3.8-Omni-Flash combines the user’s needs with the video content to identify questions worth deeper investigation, organize the key material, and search multimodal sources across the web—including images, videos, and documents. It then produces a video-centered, richly illustrated research report that helps users understand the content and solve practical problems. For example, when a user encounters color fringing around a Photoshop hair cutout, the model can break down the tutorial steps, study the principles behind Multiply and Screen blend modes, compare alternative edge-repair techniques, and explain which approach best fits the user’s situation.

Audio-Visual Production and Editing#

Qwen3.8-Omni-Flash is taking audio-visual agents into a new stage: from understanding sounds and images to independently planning, calling tools, and delivering finished videos, bringing omnimodal intelligence into professional audio-visual content production workflows.

Music2MV#

For music video (MV) creation, Qwen3.8-Omni-Flash can understand the structure, rhythm, mood, vocals, and instrumental changes of a user-provided song in fine detail, informing the design of characters, scenes, and shots. It can also output line-level lyrics with timestamps to align singing, subtitles, and visuals. Combined with creative tools such as Qwen-MM-Plugins, the model supports the complete workflow from music understanding and creative planning to final quality review, demonstrating strong audio-visual understanding, reasoning, and creation capabilities.

Workflow

Next

视频 · 前往原文观看

Final Result 1

Next

视频 · 前往原文观看

Final Result 2

Next

视频 · 前往原文观看

Final Result 3

Next

视频 · 前往原文观看

Short Drama Translation#

Traditional video translation often requires repeatedly switching between transcription, translation, dubbing, and editing platforms. This complicates API calls and workflow coordination and makes it difficult to maintain consistency across character voices, dialogue duration, and visual pacing. With an agent built on Qwen3.8-Omni-Flash, users can describe their needs in a single sentence to perform speaker-aware dialogue recognition, conversational translation, character voice cloning and dubbing, audio remixing, and final quality review. These otherwise fragmented localization steps become a complete workflow, enabling automated delivery of short dramas for international audiences.

Workflow

Next

视频 · 前往原文观看

Final Result 1

Next

Source

视频 · 前往原文观看

Translated Result

视频 · 前往原文观看

Final Result 2

Next

Source

视频 · 前往原文观看

Translated Result

视频 · 前往原文观看

Final Result 3

Next

Source

视频 · 前往原文观看

Translated Result

视频 · 前往原文观看

Long-Form Film Commentary#

Producing commentary videos for full-length films of two or three hours often requires repeatedly watching the film and reconstructing its plot, followed by shot selection, scriptwriting, voiceover, music, and editing. This is a complex and time-consuming process. With an agent built on Qwen3.8-Omni-Flash, users need only provide a film and describe their creative requirements in one sentence to perform long-form video understanding, key-plot extraction, commentary planning, voiceover and music production, editing, rendering, and final quality review. The agent can also intelligently interleave original dialogue with commentary, automatically adjusting speech rate and volume so that narration, original audio, background music, and visuals flow naturally together, creating a more authentic, immersive, and cinematic commentary video.

Workflow

Next

视频 · 前往原文观看

Final Result 1

Next

视频 · 前往原文观看

Final Result 2

Next

视频 · 前往原文观看

From Using Models to Optimizing Models#

Real-world multimodal applications are often complex and cost-sensitive, requiring models to balance quality, latency, compute, and deployment costs.

Customizing smaller models for specific scenarios is therefore an important path to deploying applications at scale. Yet traditional workflows involve data construction, problem diagnosis, multiple rounds of training, and evaluation, making them time-consuming and heavily dependent on human expertise. This time, we extend Qwen3.8-Omni-Flash into model development itself, exploring a new approach in which large models drive research and development while smaller models serve business needs.

We gave Qwen3.8-Omni-Flash a task: improve Qwen2.5-Omni-3B's Sichuan dialect speech recognition within 12 hours and deliver a usable model. It independently selected the WenetSpeech-Chuan evaluation set, fixed the evaluation criteria, and established a baseline. It then listened directly to audio samples, diagnosed problems using the recognition results, and constructed targeted training data. Across four consecutive rounds of experiments, the agent created 3,413 training examples, adjusted its approach based on evaluation feedback, retained effective improvements, and rolled back unsuccessful attempts. Qwen2.5-Omni-3B's character error rate on the same evaluation set ultimately fell from 25.79% to 15.30%, a relative reduction of approximately 40.7%.

This experiment demonstrates another possibility for model evolution: general-purpose multimodal models understand data, plan experiments, and drive iteration, while smaller models acquire specialized capabilities for specific applications. Agents can go beyond using models to help solve practical business problems.

视频 · 前往原文观看

Audio-Visual Information Compression#

Audio and video carry rich information, but their linear, unstructured form makes retrieval and reuse difficult. Qwen3.8-Omni-Flash understands content across sound, visuals, and timelines, using an agentic workflow to extract information, reorganize its structure, and verify results. It transforms the core knowledge and practical experience in long videos into denser information assets that are easier to consume and reuse.

Video2Note#

To turn video knowledge into structured resources, we have open-sourced Video2Note in Qwen-MM-Plugins. Drawing on Qwen3.8-Omni-Flash's joint understanding of speech, visuals, and procedures, it automatically organizes knowledge, breaks down key steps, selects representative frames, and generates PDF notes with corresponding text and images. Automated review and iterative correction further condense hours of video into clear, readable documents that are easy to revisit.

Workflow

Next

视频 · 前往原文观看

Final Result 1

Next

Example Video

视频 · 前往原文观看

Final PDF

Final Result 2

Next

Example Video

视频 · 前往原文观看

Final PDF

Final Result 3

Next

Example Video

视频 · 前往原文观看

Final PDF

Final Result 4

Next

Example Video

视频 · 前往原文观看

Final PDF

Omni Skill Creator#

Videos record not only "how to do something," but also the practical expertise accumulated by specialists. With this in mind, we introduce Omni Skill Creator as a new open-source capability in Qwen-MM-Plugins. It can extract standard operating procedures (SOPs) from demonstrations to perform reusable automated work, or learn tool usage, decision criteria, and key insights from expert instruction. A single demonstration becomes an agent skill that has been verified and evaluated, enabling reusable, shareable skills built from omnimodal content.

视频 · 前往原文观看

Real-Time Audio-Visual Interaction#

Qwen3.8-Omni-Flash is designed for deep understanding and creation with complete audio-visual content. For continuous, low-latency interaction, we further introduce Qwen3.8-Omni-Flash-Realtime. It perceives and responds while receiving live audio-visual streams, and uses real-time context to call tools and execute tasks, taking omnimodal capabilities from “understanding a piece of content” to “participating in an interaction.”

Real-Time Speaking Practice#

Spoken language has no standard input. Accents, vowel and consonant substitutions, and tonal deviations can cause word-for-word transcription to diverge from the intended meaning. Qwen3.8-Omni-Flash-Realtime jointly models pronunciation and semantics, understands nonstandard expressions affected by accents, aligns them with the correct words, and generates standard-pronunciation demonstrations in real time. Across multiple practice turns, the model updates its judgment with new audio, correcting errors that affect understanding while preserving natural rhythm, tone, and emotion.

视频 · 前往原文观看

Omnimodal Spatial Audio Perception#

In real-world spaces, sound provides another coordinate axis beyond vision. Qwen3.8-Omni-Flash-Realtime combines spatial sound with visual information to continuously determine the direction and distance of sound sources while perceiving obstacles, navigable areas, and changes in the scene, making it the first omnimodal model capable of locating targets by sound.

For instructions such as “come over here” or “go see what is making that sound,” the model can isolate voices and target sounds from environmental noise, ground their meaning in the surrounding space, and call tools to perform localization, search, path planning, and navigation—from hearing a target to reaching it.

视频 · 前往原文观看

External Knowledge for Audio-Visual Interaction#

Real-time interaction requires not only low latency, but also the ability to load knowledge and behavior dynamically for each application. Qwen3.8-Omni-Flash-Realtime supports injecting identity settings, expression styles, business knowledge, and interaction rules through Skills, while tool use extends these capabilities into task execution.

In scenarios such as customer service, the model can load brand language and service procedures in real time, understand the user's speech, visuals, and context, generate responses that follow business requirements, and execute actions. The same real-time model can therefore take on different knowledge, roles, and ways of acting.

视频 · 前往原文观看

Benchmark Results#

Omni#

Qwen3.8-Omni-FlashQwen3.5-Omni-PlusGemini 3.8
Flash
Seed 2.0 LiteMuse Spark 1.2
Agentic Omni Intelligence

WildClawBench-MM

Multimodal tool use

71.034.558.941.9--

UniClawBench

Multimodal tool use

69.667.169.061.2--

AgenticVBench

Multimodal tool use

36.814.545.010.0--

OmniGAIA

Web Search

74.057.278.664.4--
General Audio-Visual Capabilities

DailyOmni

Audio-Visual Understanding

85.185.184.081.479.6

WorldSense

Audio-Visual Understanding

68.563.969.667.365.0

AVUT

Audio-Visual Understanding

86.685.988.081.582.4

JoinAVBench

Audio-Visual Understanding

75.974.170.470.671.8

OmniVideoBench

Audio-Visual Reasoning

63.453.865.258.562.2

Video-MME-v2

Audio-Visual Reasoning

65.047.971.064.9--

LVOmniBench

Long Video Reasoning

63.353.270.7----

OmniCloze

Audio-Visual Caption

63.264.260.956.365.3

OmniCap-IF

Audio-Visual Caption

CSR

80.6

ISR

28.2

CSR

72.1

ISR

14.1

CSR

81.9

ISR

28.3

CSR

74.6

ISR

18.1

CSR

77.9

ISR

26.8

QIVD

Audio-Visual Interaction

69.665.669.162.062.0

StreamingBench

Audio-Visual Interaction

80.857.179.977.277.8
General Audio Capabilities

AliMeeting Test

Multi-Speaker ASR (DER | cpWER, ↓)

3.4 | 17.288.1 | 89.672.6 | 53.175.1 | 76.193.7 | 92.7

AISHELL-4

Multi-Speaker ASR (DER | cpWER, ↓)

2.8 | 11.2100.0 | 100.066.4 | 56.964.8 | 64.291.3 | 86.0

MagicData-RAMC

Multi-Speaker ASR (DER | cpWER, ↓)

5.7 | 14.198.4 | 97.167.9 | 33.843.4 | 35.182.1 | 75.3

MLC-SLM (en)

Multi-Speaker ASR (DER | cpWER, ↓)

4.0 | 14.268.6 | 63.960.8 | 26.640.4 | 45.574.3 | 52.9

WenetSpeech (Net)

ASR (WER, ↓)

4.83.714.24.368.2

WenetSpeech (Meeting)

ASR (WER, ↓)

4.64.816.74.742.6

FLEURS-ASR

Multilingual ASR (WER, ↓)

9.37.27.932.123.6

FLEURS-S2TT

Multilingual S2TT (BLEU)

31.832.233.024.828.8

SpotSoundBench

Audio Grounding

67.264.239.759.616.9

MMAU

Audio Understanding

81.881.976.977.263.5

MMAR

Audio Understanding

79.879.878.577.767.3

MMSU

Audio Understanding

82.183.083.380.259.9

LongAudioSpan

Long Audio Reasoning

Accuracy

82.7

Rubric

71.8

Chain

48.2

Accuracy

74.4

Rubric

49.8

Chain

45.1

Accuracy

79.3

Rubric

65.5

Chain

64.6

----

MuchoMusic-RUL

Music Understanding

72.671.653.761.740.1

HumMusQA

Music Understanding

75.875.571.266.063.3

MusTBench

Music Understanding

50.649.140.344.029.4

Audio MultiChallenge

Audio Interaction

71.557.671.963.457.9

WildSpeech

Audio Interaction

74.375.776.474.573.4

VoiceBench

Audio Interaction

91.692.992.384.179.8

1. Evaluation harnesses for Agentic Omni Intelligence: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, and OmniGAIA uses no harness. For WildClawBench-MM, we evaluate only the multimodal tasks in WildClawBench that involve images, video, or audio.
2. FLEURS: ASR and S2TT evaluation results both cover the following 60 languages: Chinese (Mandarin), English, Cantonese, Arabic, German, French, Spanish, Portuguese, Indonesian, Italian, Korean, Russian, Thai, Vietnamese, Japanese, Turkish, Hindi, Malay, Dutch, Urdu, Norwegian, Swedish, Danish, Hebrew, Finnish, Polish, Icelandic, Czech, Filipino, Persian, Greek, Afrikaans, Asturian, Belarusian, Bulgarian, Bengali, Bosnian, Catalan, Cebuano, Estonian, Galician, Gujarati, Croatian, Hungarian, Javanese, Kazakh, Kannada, Kyrgyz, Latvian, Macedonian, Malayalam, Marathi, Punjabi, Romanian, Slovak, Slovenian, Swahili, Tajik, Azerbaijani, and Ukrainian.
3. Empty cells (--): scores are not yet available or are not applicable.

Agentic Omni Understanding#

Key information in long audio and video recordings is often scattered across different segments, while complex questions require multiple steps of reasoning across sound and images. Agentic Omni Understanding enables the model to start from the question, plan its approach, call tools, and progressively locate and verify evidence. By focusing computation on relevant content, it improves the accuracy and efficiency of long-form audio-visual understanding. To evaluate this capability, we compare Qwen3.8-Omni-Flash and Gemini 3.8 Flash on OmniVideoBench, Video-MME-v2, and LVOmniBench under two settings: Static, where the model directly interprets the input, and an agent mode using Qwen Code. This comparison shows how introducing agent workflows affects each model’s performance.

Qwen3.8-Omni-Flash
(Static)
Qwen3.8-Omni-Flash
(Qwen Code)
Gemini 3.8 Flash
(Static)
Gemini 3.8 Flash
(Qwen Code)

OmniVideoBench

Audio-Visual Reasoning

63.467.865.270.1

Video-MME-v2

Audio-Visual Reasoning

65.071.371.072.7

LVOmniBench

Long Video Reasoning

63.373.670.770.7

Text#

Qwen3.8-Omni-FlashQwen3.8-
Flash
Qwen3.8-
27B
Qwen3.7-
Plus
DeepSeek-V4-Flash-0731Claude-Opus-4.6 (Max)
Coding and Agent

DeepSWE 1.1

Long-horizon software engineering

57.858.742.216.554.4--

SWE-bench Pro

Long-horizon software engineering

63.362.561.755.856.053.4

SWE-bench Multilingual

Multilingual software engineering

80.581.073.875.8--77.5

NL2Repo-Bench

Repo-level code generation

48.948.142.341.154.247.6

CoWorkBench

Long-horizon office work

75.373.970.765.145.168.2
General Text Capabilities

IFBench

Instruction following

81.581.379.579.179.262.5

GPQA Diamond

Scientific reasoning

91.091.789.290.390.891.3

HLE

Multidisciplinary reasoning

36.535.930.834.733.840.0

LiveCodeBench v6

Competitive coding

92.691.990.389.690.688.8

1. DeepSWE 1.1: evaluated with the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, 256K context window. We report the highest score across the two harnesses; notably, Qwen3.8-Flash-Next performs best on mini-SWE-agent.
2. SWE-bench Pro: except for Claude-Opus-4.6 (Max), for which we report the officially published score, all models are evaluated with the Claude Code harness, temp=1.0, top_p=0.95, 256K context window. Problematic tasks were corrected and all baseline models were re-evaluated on the refined benchmark.
3. SWE-bench Multilingual: evaluated with the mini-SWE-agent harness, temp=1.0, top_p=0.95, 256K context window.
4. NL2Repo-Bench: evaluated with the Claude Code harness. To prevent reward hacking, we disable Bash commands that attempt to access the specific repository, such as pip download, pip install and git clone.
5. CoWorkBench: an in-house cowork benchmark for evaluating long-horizon office and productivity agent tasks across computer science, finance, law, medical and other productivity domains.
6. HLE: judged by GPT-4o.
7. Empty cells (--): scores are not yet available or are not applicable.

Vision#

Qwen3.8-Omni-FlashQwen3.8-
Flash
Qwen3.8-
27B
Qwen3.7-
Plus
Claude-Opus-4.6 (Max)
Agentic Vision Intelligence

ClawEval-MM

Multimodal tool use

Pass@3

60.4

Average

61.9

Pass@3

64.4

Average

60.4

Pass@3

57.4

Average

56.9

Pass@3

57.4

Average

60.1

Pass@3

52.5

Average

54.7

AndroidWorld

Mobile use

87.184.581.981.062.0

Vision2Web

Visual web development

62.964.062.942.1--
General Vision Capabilities

ERQA

Embodied intelligence

71.072.365.569.840.8

LVBench

Long video understanding

76.976.672.476.263.0

RealWorldQA

Real-world perception

87.788.585.986.973.9

MathVision

Visual math problem solving

Without CI

91.8

With CI

96.2

Without CI

90.6

With CI

95.7

Without CI

90.0

With CI

94.6

Without CI

90.3

With CI

88.4

Without CI

65.5

With CI

--

CharXiv (RQ)

Scientific chart analysis

Without CI

83.5

With CI

91.4

Without CI

84.6

With CI

90.6

Without CI

83.7

With CI

90.2

Without CI

85.8

With CI

85.9

Without CI

66.0

With CI

--

1. ClawEval-MM: scores are reported as "pass@3 / average score". Pass@3 measures the percentage passed in at least one of three trials, and the average score is the mean score across the three trials.
2. RecreationBench: an in-house long-horizon application-recreation benchmark for evaluating hybrid-agent abilities spanning five platforms — desktop (Ubuntu, macOS, Windows), mobile (Android) and web.
3. OSWorld 2.0: scores are reported as "binary / partial". The binary score is the percentage of tasks that receive the full task reward, while the partial score aggregates the partial rewards obtained across all tasks.
4. Vision2Web: scores are reported as the average over the frontend, webpage and website categories, using the Claude Code harness and judged by gpt-5.4-2026-03-05.
5. MathVision, CharXiv (RQ): scores are reported as "without CI / with CI". A small number of incorrect ground-truth annotations in MathVision were corrected after manual verification. Our model's score is evaluated using a fixed prompt, e.g. "Please reason step by step, and put your final answer within \boxed{}." For other models, we report the higher score between runs with and without the \boxed{} formatting.
6. Empty cells (--) indicate scores not yet available or not applicable.

Throughput & Latency#

The following results show the observed throughput and latency of the Qwen3.8-Omni-Flash-Realtime API under different input conditions, reflecting the performance users experience in production environments.

Input ScenarioText Output TPS (Tokens/s)Time to First
Token (ms)
Time to First Audio Packet (ms)Audio Generation RTF
Realtime API Performance
Audio 6s84.87591.26978.360.1538
Audio 12s83.18604.80982.740.1537
Audio 20s81.06617.981026.390.1538
Audio-Visual 6s84.89837.961214.730.1524
Audio-Visual 12s84.33911.851268.070.1527
Audio-Visual 20s83.00981.011350.490.1528

Supported Languages#

CapabilityLanguagesChinese Dialects
Speech Recognition74 languages: Afrikaans, Arabic, Asturian, Azerbaijani, Basque, Belarusian, Bengali, Bosnian, Bulgarian, Cantonese, Catalan, Cebuano, Chinese, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Filipino, Finnish, French, Galician, Georgian, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Interlingua, Italian, Japanese, Javanese, Kannada, Kazakh, Korean, Kyrgyz, Lingala, Latvian, Lithuanian, Macedonian, Malay, Malayalam, Maltese, Maori, Marathi, Mongolian, Norwegian Bokmål, Norwegian Nynorsk, Oriya, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tajiki, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Uyghur, and Vietnamese39 dialects: Northeastern Mandarin, Guizhou dialect, Guangdong Cantonese, Henan dialect, Hong Kong Cantonese, Shanghainese, Shaanxi dialect, Tianjin dialect, Taiwanese Mandarin, Yunnan dialect, Anhui dialect, Fujian dialect, Gansu dialect, Guangdong Mandarin, Hubei dialect, Hunan dialect, Jiangxi dialect, Shandong dialect, Shanxi dialect, Sichuanese, Guangxi dialect, Hainan dialect, Chongqing dialect, Changsha dialect, Hangzhou dialect, Hefei dialect, Yinchuan dialect, Zhengzhou dialect, Shenyang dialect, Wenzhou dialect, Wuhan dialect, Kunming dialect, Taiyuan dialect, Nanchang dialect, Jinan dialect, Lanzhou dialect, Nanjing dialect, Hakka, and Southern Min
Speech Generation29 languages: Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian, Thai, Indonesian, Arabic, Vietnamese, Turkish, Finnish, Polish, Hindi, Dutch, Czech, Urdu, Tagalog, Swedish, Danish, Hebrew, Icelandic, Malay, Norwegian, and Persian7 dialects: Sichuanese, Beijing dialect, Tianjin dialect, Nanjing dialect, Shaanxi dialect, Cantonese, and Southern Min

Getting Started with Qwen3.8-Omni-Flash#

API Usage#

Qwen3.8-Omni-Flash officially supports reasoning_effort to adjust reasoning depth and control costs:

  • xhigh (default): for complex tasks that require in-depth analysis.
  • medium: balances accuracy and speed.
  • low: efficient reasoning optimized for speed and cost.

In addition, preserve_thinking is enabled by default across all scenarios for the best out-of-the-box experience.

Qwen3.8-Omni-Flash supports industry-standard protocols, including OpenAI-compatible Chat Completions and Responses APIs. Examples follow:

"""
Environment variables:
  DASHSCOPE_API_KEY: Your API Key from https://platform.qianwenai.com/home
  DASHSCOPE_BASE_URL: (optional) Base URL for compatible-mode API.
    - Beijing: https://dashscope.aliyuncs.com/compatible-mode/v1
    - Singapore: https://dashscope-intl.aliyuncs.com/compatible-mode/v1
"""
from openai import OpenAI
import os
api_key = os.environ.get("DASHSCOPE_API_KEY")
if not api_key:
    raise ValueError(
        "DASHSCOPE_API_KEY is required. "
        "Set it via: export DASHSCOPE_API_KEY='your-api-key'"
    )
client = OpenAI(
    api_key=api_key,
    base_url=os.environ.get(
        "DASHSCOPE_BASE_URL",
        "https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
    ),
)
messages=[
    {
        "role": "user",
        "content": [
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241022/emyrja/dog_and_girl.jpeg"
                },
            },
            {
                "type": "input_audio",
                "input_audio": {
                    "data": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250211/tixcef/cherry.wav",
                    "format": "wav"
                },
            },
            {"type": "text", "text": "Please describe the image and tell me what is being said in the audio."},
        ],
    },
]
completion = client.chat.completions.create(
    model="qwen3.8-omni-flash",
    messages=messages,
    extra_body={
        "enable_thinking": True,
        # "preserve_thinking": True,
    },
    reasoning_effort="xhigh",  # supported levels are xhigh, medium, and low
    stream=True,
)
reasoning_content = ""
answer_content = ""
is_answering = False
print("\n" + "=" * 20 + "Reasoning" + "=" * 20 + "\n")
for chunk in completion:
    if not chunk.choices:
        print("\nUsage:")
        print(chunk.usage)
        continue
    delta = chunk.choices[0].delta
    if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None:
        if not is_answering:
            print(delta.reasoning_content, end="", flush=True)
        reasoning_content += delta.reasoning_content
    if hasattr(delta, "content") and delta.content:
        if not is_answering:
            print("\n" + "=" * 20 + "Answer" + "=" * 20 + "\n")
            is_answering = True
        print(delta.content, end="", flush=True)
        answer_content += delta.content

Best Practices for Controllable Audio-Visual Captioning#

Content organization and level of detail: Organize video descriptions, OCR extraction, audio descriptions, and speech transcripts into sections covering visuals, speech, music, sound effects, and ambient sounds. Preserve the original text, speaker identities, and corresponding time ranges to make retrieval and verification easier. Specify a target length in the prompt to control detail, for example, Describe the video in approximately 2000–3000 words., and adjust it to the video’s duration and information density.

Prompt Example

Provide a detailed description of the video.
Make sure your description covers every one of the following dimensions:
Visual
- Subjects and characters: appearance, clothing, gender/age cues, identity, distinctive features
- Actions and events in chronological order, and how the scene evolves over time
- Setting and background: location, environment, time of day
- Spatial layout and relations between subjects/objects; counts and quantities
- On-screen text: captions, titles, subtitles, logos, UI — exact content and appearance
- Visual style: colors, lighting, camera shots, angles, and camera movement
Audio
- Speech: the exact spoken content, transcribed verbatim
- Speakers: who is speaking (mapped to the on-screen person or voice-over), with accent, tone, gender/age cues
- Speaking state: prosody, emotion, volume, and speaking style
- Music: presence, genre/mood, and lyrics if any
- Sound effects and ambient/background sounds
- Non-speech vocalizations: laughter, crying, applause, etc.
Audio-visual correspondence
- Which speech or sound aligns with which on-screen person or visual event
- The timing of each event, expressed with timestamps
It should explicitly include three sections:
1. A structured chronological storyline of **every noticeable audio and visual details**
2. A structured list of all visible text. For each text element, include start timestamp, end timestamp, the exact text content, the appearance characteristics. If no text appears, explicitly state so.
3. A structured speech-to-text transcription, include speaker (corresponding to the character or voice-over in Section 1, including their accent and tone), exact spoken content, start timestamp, end timestamp, and speaking state (prosody, emotion, and style). If no speech appears, explicitly state so.
Aside from these three required sections, you are free to organize any additional content in any way you find helpful. This additional content can include global information about the entire video or localized information about specific moments. You may choose the topic of this extra content freely.
Rules:
- Add as much descriptive detail as possible.
- Do not use Markdown bold formatting.
- Carefully look at frames and listen to the audio, making sure no detail is overlooked.
Output Format:
```
## Storyline
<xx:xx.xxx> - <xx:xx.xxx>
<an unstructured long paragraph in natural language describing what happened during this period, blending both audio and video details.>
<xx:xx.xxx> - <xx:xx.xxx>
<an unstructured long paragraph in natural language describing what happened during this period, blending both audio and video details.>
<xx:xx.xxx> - <xx:xx.xxx>
<an unstructured long paragraph in natural language describing what happened during this period, blending both audio and video details.>
...
## Visible Text
<xx:xx.xxx> - <xx:xx.xxx>
“<element>”: <appearance>
“<element>”: <appearance>
<xx:xx.xxx> - <xx:xx.xxx>
“<element>”: <appearance>
“<element>”: <appearance>
“<element>”: <appearance>
<xx:xx.xxx> - <xx:xx.xxx>
“<element>”: <appearance>
...
## Speakers and Transcript
Speaker profiles:
<speaker> - <profile>
<speaker> - <profile>
<speaker> - <profile>
...
<xx:xx.xxx> - <xx:xx.xxx>
Speaker: <speaker>
State: <description>
Content: “<content>”
<xx:xx.xxx> - <xx:xx.xxx>
Speaker: <speaker>
State: <description>
Content: “<content>”
<xx:xx.xxx> - <xx:xx.xxx>
Speaker: <speaker>
State: <description>
Content: “<content>”
...
## <another section>
<paragraphs>
## <another section>
<paragraphs>
...
```

Structured output and schema constraints: Include task instructions and a complete JSON Schema in the prompt, specifying field meanings, types, required fields, and constraints to support programmatic parsing and downstream use.

Prompt Example

Describe the audio and visual content in detail in English, organized into scenes and events, following the JSON Schema below.
All timestamps must be relative to the beginning of the video. End times must not precede start times or exceed the video duration. Each event must fall within the time range of its parent scene.
Include only information directly supported by the audio or video. Do not guess or invent details. Do not infer causality merely because a sound and an action occur at the same time.
Return only valid JSON, without Markdown fences or commentary.
JSON Schema:
{
  "$defs": {
    "Event": {
      "additionalProperties": false,
      "properties": {
        "time_range": {
          "$ref": "#/$defs/TimeRange",
          "description": "Time range of the event"
        },
        "participants": {
          "description": "People, animals, or objects involved, named by observable features; use consistent names for the same participant",
          "items": {
            "type": "string"
          },
          "title": "Participants",
          "type": "array"
        },
        "action": {
          "description": "Specific actions, interactions, and observable outcomes",
          "title": "Action",
          "type": "string"
        },
        "sounds": {
          "description": "Sounds heard during the event; use an empty list if none are discernible",
          "items": {
            "type": "string"
          },
          "title": "Sounds",
          "type": "array"
        }
      },
      "required": [
        "time_range",
        "participants",
        "action",
        "sounds"
      ],
      "title": "Event",
      "type": "object"
    },
    "Scene": {
      "additionalProperties": false,
      "properties": {
        "time_range": {
          "$ref": "#/$defs/TimeRange",
          "description": "Time range of the scene"
        },
        "setting": {
          "description": "Environment, spatial layout, and main visual features",
          "title": "Setting",
          "type": "string"
        },
        "events": {
          "description": "Events in chronological order; use an empty list if there are none",
          "items": {
            "$ref": "#/$defs/Event"
          },
          "title": "Events",
          "type": "array"
        }
      },
      "required": [
        "time_range",
        "setting",
        "events"
      ],
      "title": "Scene",
      "type": "object"
    },
    "TimeRange": {
      "additionalProperties": false,
      "properties": {
        "start_seconds": {
          "description": "Start time in seconds relative to the beginning of the video",
          "minimum": 0,
          "title": "Start Seconds",
          "type": "number"
        },
        "end_seconds": {
          "description": "End time in seconds; must not precede the start time",
          "minimum": 0,
          "title": "End Seconds",
          "type": "number"
        }
      },
      "required": [
        "start_seconds",
        "end_seconds"
      ],
      "title": "TimeRange",
      "type": "object"
    }
  },
  "additionalProperties": false,
  "properties": {
    "summary": {
      "description": "An overview of the main content of the video",
      "title": "Summary",
      "type": "string"
    },
    "scenes": {
      "description": "Scenes in chronological order; group continuous footage with a consistent setting into one scene",
      "items": {
        "$ref": "#/$defs/Scene"
      },
      "title": "Scenes",
      "type": "array"
    }
  },
  "required": [
    "summary",
    "scenes"
  ],
  "title": "CaptionResult",
  "type": "object"
}

Installing Qwen-MM-Plugins in an Agent Harness#

Qwen-MM-Plugins is a multimodal plugin suite for agent harnesses. It gives agents the ability to understand images, audio, video, and documents, as well as maintain memory for long-form video and create content. It supports agent harnesses including Codex, Claude Code, Qwen Code, Gemini CLI, Qoder, CodeBuddy, and OpenClaw. We also welcome contributions from community developers.

You can enter the following request directly in your usual office agent:

Help me install the core, api, and omni-related plugins from https://github.com/QwenLM/Qwen-MM-Plugins.
Installing Qwen-MM-Plugins in an office agent using natural language

Alternatively, install from the command line:

# Run the official guided installer:
curl -fsSL https://raw.githubusercontent.com/QwenLM/Qwen-MM-Plugins/main/install.sh | bash

In the menu, select:

  • Install
  • The agent harness you use
  • The Omni plugins you need
PluginCapability
coreRead images, video frames, PDFs, Office documents, code, data, and 3D files
apiImage understanding, OCR, object localization, audio-visual transcription, speaker diarization, event analysis, and image segmentation
omni-chatcutCreate MVs, long-form film commentary, and video speech translations
omni-video2noteTurn video tutorials into PDF notes with key screenshots
omni-skill-creatorTurn instructional videos, screen recordings, or operation demonstrations into reusable Agent Skill.md files
omni-memoryBuild memory for people, dialogue, sounds, and events in long-form videos

After installation, restart the agent harness or create a new task.

@song.mp3 Generate a complete MV based on the song's rhythm and content.
@short_drama.mp4 Translate the video into English while preserving the original speakers' voice characteristics where possible.
@movie.mp4 Create a film commentary video with Chinese narration and subtitles.
@weekly_report_sop.mp4 Turn this screen recording of writing a weekly report into a Skill.md file.
@tutorial.mp4 Turn the tutorial into PDF notes with key screenshots, timestamps, and step-by-step instructions.
@documentary.mp4 Build audio-visual memory that records people, dialogue, sounds, and important events.

Getting Started with Qwen3.8-Omni-Flash-Realtime#

API Usage#

Qwen3.8-Omni-Flash-Realtime supports connections over WebSocket and WebRTC. Running the basic examples below opens the camera and microphone for a real-time audio-visual conversation. We recommend using headphones.

# Run pip install websocket-client pyaudio dashscope opencv-python -U to install dependencies
"""
Environment variables:
  DASHSCOPE_API_KEY: Your API Key from https://platform.qianwenai.com/home
  DASHSCOPE_BASE_URL: (optional) Base URL for realtime API.
    - Beijing: wss://dashscope.aliyuncs.com/api-ws/v1/realtime
    - Singapore: wss://dashscope-intl.aliyuncs.com/api-ws/v1/realtime
"""
import os
import base64
import time
import pyaudio
import cv2
from dashscope.audio.qwen_omni import MultiModality, AudioFormat,OmniRealtimeCallback,OmniRealtimeConversation
import dashscope
url = f'wss://dashscope-intl.aliyuncs.com/api-ws/v1/realtime'
dashscope.api_key = os.getenv('DASHSCOPE_API_KEY')
# Determine the voice
voice = 'Tina'
# Determine the model
model = 'qwen3.8-omni-flash-realtime'
# Determine the model role
instructions = "You are Qwen-Omni, a helpful assistant."
video_fps = 1
video_size = (1280, 720)
class SimpleCallback(OmniRealtimeCallback):
    def __init__(self, pya):
        self.pya = pya
        self.out = None
    def on_open(self):
        # Initialize audio output stream
        self.out = self.pya.open(
            format=pyaudio.paInt16,
            channels=1,
            rate=24000,
            output=True
        )
    def on_event(self, response):
        if response['type'] == 'response.audio.delta':
            # Play audio
            self.out.write(base64.b64decode(response['delta']))
        elif response['type'] == 'conversation.item.input_audio_transcription.delta':
            # Streaming preview: text is the confirmed prefix, stash is the confirmed suffix
            preview = response.get('text', '') + response.get('stash', '')
            print(f"\r[User] {preview}", end='', flush=True)
        elif response['type'] == 'conversation.item.input_audio_transcription.completed':
            # Transcription completed, print the final text and a new line
            print(f"\r[User] {response['transcript']}")
        elif response['type'] == 'response.audio_transcript.done':
            # Print the assistant's response text
            print(f"[LLM] {response['transcript']}")
# 1. Initialize audio device
pya = pyaudio.PyAudio()
# 2. Create callback function and session
callback = SimpleCallback(pya)
conv = OmniRealtimeConversation(model=model, callback=callback, url=url)
# 3. Establish connection and configure session
conv.connect()
conv.update_session(output_modalities=[MultiModality.AUDIO, MultiModality.TEXT], voice=voice, instructions=instructions)
# 4. Initialize audio input stream
mic = pya.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True)
camera = cv2.VideoCapture(0)
next_frame_at = 0.0
# 5. Main loop to process audio and video input
try:
    if not camera.isOpened():
        raise RuntimeError("Cannot open camera 0.")
    camera.set(cv2.CAP_PROP_FRAME_WIDTH, video_size[0])
    camera.set(cv2.CAP_PROP_FRAME_HEIGHT, video_size[1])
    camera.set(cv2., 1)
    print(f"Conversation started with camera ({video_fps} fps, {video_size[0]}x{video_size[1]}), speak into the microphone (Ctrl+C to exit)...")
    while True:
        audio_data = mic.read(3200, exception_on_overflow=False)
        conv.append_audio(base64.b64encode(audio_data).decode())
        success, frame = camera.read()
        if not success:
            raise RuntimeError("Cannot read a camera frame.")
        if time.monotonic() >= next_frame_at:
            frame = cv2.resize(frame, video_size)
            success, image = cv2.imencode('.jpg', frame, [cv2.IMWRITE_JPEG_QUALITY, 80])
            if not success:
                raise RuntimeError("Cannot encode a camera frame.")
            conv.append_video(base64.b64encode(image).decode())
            next_frame_at = time.monotonic() + 1 / video_fps
        time.sleep(0.01)
except KeyboardInterrupt:
    pass
finally:
    # Clean up resources
    camera.release()
    conv.close()
    mic.close()
    if callback.out:
        callback.out.close()
    pya.terminate()
    print("\nConversation ended")

Qwen-Live Harness#

Qwen-Live Harness is a comprehensive open-source harness designed around the Qwen3.8-Omni-Flash-Realtime API. It can be installed with a single command and integrated into mainstream agent workflows. It supports task delegation, proactive interaction, long-term memory, and context management, and welcomes contributions from the community.

Qwen-Live Harness interaction framework

Figure 2. Qwen-Live Harness Interaction Framework.

Install and get started:

npm install -g qwen-live-harness
qwen-live-harness init
qwen-live-harness

Citation#

@misc{qwen38omniflash,
    title = {Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.},
    url = {https://qwen.ai/blog?id=qwen3.8-omni-flash},
    author = {{Qwen Team}},
    month = {September},
    year = {2026}
}

来源:Qwen:Blog Retrieval(API) · qwen.ai