Last updated: September 2026 — breaking coverage
Imagine handing an AI a 90-minute podcast recording, a folder of raw event footage, or an entire season of short videos — and getting back speaker-labelled transcripts, a highlight plan, subtitled clips, and even a dubbed version, in one pipeline. That is the bet behind Qwen3.8-Omni-Flash, the native omni-modal model Alibaba's Qwen team launched on September 18, 2026, now available on the Qianwen AI Platform with a 1M-token context window across text, image, audio, and video input.
The official announcement, titled "Omni Senses. Agentic Delivery.", is explicit about the ambition: moving omni-modal models from understanding content to planning tasks, calling tools, and completing creative work. For anyone who produces content for a living, that framing — not the benchmark table — is the story.
The numbers that matter
- 98%+ cut in the API price per hour of audio input and 93%+ for audio-visual input versus the previous generation (Qwen3.5-Omni-Plus), per Qwen's documented pricing methodology (text prices in CNY per 1M tokens; AV measured at 720p, 1 fps).
- +25% average improvement across 29 evaluations over Qwen3.5-Omni-Plus, with agentic audio-visual gains of 36.5 points on WildClawBench-MM, 22.3 points on AgenticVBench, and a 69.6 on UniClawBench.
- Long-form efficiency: agentic understanding on OmniVideoBench raised accuracy from 63.4 to 67.8 while cutting tokens per query from 145,736 to 79,117 — roughly 45.7% cheaper inference for the same question about a long video.
- Speech accuracy: AliMeeting DER dropped from 88.11 to 3.35, and meeting transcription quality (cpWER) improved from 89.61 to 17.18.
- Competitive claim: audio-visual performance "close to Gemini 3.8 Flash" and overall audio performance "exceeding Gemini 3.8 Flash" — vendor-reported, so treat it as a claim, not a verdict.

Five workflows it changes for creators
- Single-sentence localization. The announcement describes an agent that handles speaker-aware dialogue recognition, conversational translation, character voice cloning and dubbing, audio remixing, and final quality review — the fragmented localization stack collapsing into one described workflow.
- Music video production. The model outputs line-level lyrics with timestamps to align singing, subtitles, and visuals, and plans the creative work from music understanding through final review.
- Video editing and film commentary, named directly among the target agentic applications, alongside audio-visual summarization and video-centred deep research reports.
- Controllable captioning. You define the subject, time range, level of detail, and output format — an overview for publishing, segment-level retrieval for editing, or structured analysis of shots, lighting, and sound for an archive.
- Meetings and live interaction. Up to one hour of native audio-visual input with speaker segmentation for multi-participant meetings, plus a Qwen3.8-Omni-Flash-Realtime variant that perceives and responds during live streams — including accent-tolerant speaking practice and spatial audio perception Qwen calls a first for omni-modal models.

How to access it today
The model is served through the Qianwen AI Platform at qwen.ai with two endpoints — Qwen3.8-Omni-Flash and Qwen3.8-Omni-Flash-Realtime. Qwen Studio apps exist for web, iOS, Android, macOS, and Windows if you want to test capabilities before building. For engineers, the team open-sourced Qwen-Live Harness (installable with npm install -g qwen-live-harness), a native runtime for continuous real-time omni-modal interaction, and expanded Qwen-MM-Plugins, which brings on-demand perception, tool use, and workflow execution to long-form audio and video.
For broader context on the Qwen family this season, see our coverage of the Qwen3.8-Flash text model and its million-token context.
Quick comparison
| Aspect | Qwen3.8-Omni-Flash | Qwen3.5-Omni-Plus | Gemini 3.8 Flash |
|---|---|---|---|
| Input modalities | Text, image, audio, video (native omni) | Omni-modal | Multi-modal |
| Context window | 1M tokens | Smaller | Large |
| Price per hour of audio input | >98% lower than predecessor | Baseline | Compared in Qwen's official figure |
| Agentic AV tasks | +36.5 pts WildClawBench-MM | Baseline | "Close" per Qwen's claim |
| Realtime variant | Yes (Omni-Flash-Realtime) | No | Yes (Live family) |
Why this launch matters now
Omni-Flash lands as the third major Qwen-family update in under two months, and it marks a clear shift in how model vendors compete for content work in 2026: the fight has moved from "who answers more accurately" to "who executes more affordably". When the price of an hour of audio input collapses by more than 98% in a single generation, tasks that used to be agency services — podcast archive analysis, webinar summarization, short-form localization — become line items a solo creator can expense weekly.
There is also a strategic signal here for anyone building on these APIs. Qwen paired the model with open tooling (Qwen-Live Harness, Qwen-MM-Plugins) instead of keeping the whole stack closed, which suggests where it expects developers to build: harnesses, plugins, and vertical workflows on top of commodity-priced omni perception. Google, OpenAI, and Meta are all pushing toward agents that see and hear; Qwen's differentiator this round is pairing the narrative with published numbers and a runtime you can install today. Watch two things over the coming weeks: independent benchmark runs on the audio claims, and whether an open-weights Omni release follows — both would change the build-versus-buy calculus for production teams.
A practical way to size the opportunity: list every recurring task in your production pipeline that starts with "watch this footage" or "listen to this recording" — episode transcripts, meeting-to-content summaries, clip mining from long streams, first-pass subtitle drafts, translated dubs for a second-market channel. Price each one at your current tool or vendor rate, then re-price it at the published Omni-Flash rates. For most teams doing weekly long-form output, audio and video ingestion turns out to be a third or more of the content budget, which is exactly the slice this launch attacks. Even if you don't switch tools this quarter, that arithmetic is worth having in hand before your next renewal or hire decision.
Honest limitations
- No open weights (yet). Unlike its siblings Qwen3.8-27B and Flash-Next on Hugging Face, Omni-Flash had no public weights at launch — you run it on Alibaba's cloud, under its pricing and its data handling.
- Vendor benchmarks only. Every comparison above, including the Gemini claims, comes from Qwen's own announcement and has not been independently verified.
- CNY pricing with a specific methodology (720p at 1 fps for AV estimates) — model your own workload before committing a production pipeline to it.
- Multilingual speech was benchmarked, not promised. The evaluation suite includes FLEURS-60 ASR and speech-to-text translation (60 languages), but the announcement makes no product-level promises about any specific language's quality.
- Voice cloning deserves governance. A single-sentence dubbing workflow is powerful and risky at once; human review and clear rights to cloned voices are on you.
By the way — if your pipeline also needs writing and editing in Arabic, ArWriter covers that end of the workflow with a full Arabic-first interface, from drafting to polish.
Frequently Asked Questions
What is Qwen3.8-Omni-Flash?
It is Alibaba's native omni-modal AI model launched on September 18, 2026, accepting text, image, audio, and video with a 1M-token context, and designed to execute complete content-production workflows — editing, dubbing, captioning, summarization — rather than only understand them.
Is Qwen3.8-Omni-Flash free to use?
No. It is available through the Qianwen AI Platform on consumption-based pricing that Qwen says cuts the cost of an hour of audio input by more than 98% versus the previous generation, with text prices quoted in CNY per million tokens.
Can it run locally?
Not the model itself — no open weights were released for Omni-Flash at launch. What is open source is the surrounding tooling: Qwen-Live Harness and the expanded Qwen-MM-Plugins.
How does it compare with Gemini 3.8 Flash?
Per Qwen's own announcement, audio-visual performance is close to Gemini 3.8 Flash while overall audio performance exceeds it. These are vendor-run comparisons, so independent validation is still pending.
Who should care most?
Creators and teams producing podcasts, video series, e-learning, and localized content, where long audio-video inputs and translation or dubbing steps dominate cost and time.
Sources
- Qwen official announcement: Qwen3.8-Omni-Flash — Omni Senses. Agentic Delivery. — primary source for all figures (September 18, 2026)
- Qwen / Qianwen AI Platform — availability, API endpoints, and Qwen Studio apps
- ArWriter's earlier coverage of Qwen3.8-Flash and the Qwen 3.8-Max GA — family context
And if your production chain ends where the AI model's output begins — writing, refining, and publishing in Arabic — ArWriter's toolset picks it up from there.