Alibaba releases Qwen 3.8 Omni Flash multimodal model
New cloud-only omni model handles text, image, audio, and video inputs with text and speech outputs.
What to know
- Qwen 3.8 Omni Flash is a closed-source, cloud-only multimodal model released September 14 that processes text, image, audio, and video inputs.
- The model is available exclusively via Alibaba Cloud APIs across five geographic regions, with unified token pricing starting at $0.15 per million input tokens.
- Only companion repositories are open-sourced; the model itself and its architecture remain proprietary, with community criticism focused on the lack of open weights.
The closed-source, cloud-only approach limits accessibility compared to open alternatives.
-
“No OSS :/”
Community · tokenstead.ai ↗
Alibaba Model developerQwen team Development and release team
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Community reacts to closed-source model
Early community feedback noted the lack of open-source weights as a significant limitation, with responses characterizing the release as proprietary.
“No OSS :/…”
— tokenstead.ai community reaction -
🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities! Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result. Highlights: 🥳 -
2 more of the top 3 · 21 posts in this stretch
-
I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.E.g. I…
-
I use models.dev's CLI tool, which I think gets data from OpenRouter, and ArtificialAnalysis so your coding agent can help you narrow it down.<sidenote>Similarly, HuggingFace has a CLI + a few skills, and they are very useful.I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20…
-
-
1
Technical details surface: pricing and API structure disclosed
Technical details reveal the model is API-only on Alibaba Cloud endpoints (Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, US-Virginia), with unified token billing at $0.15/1M input on Singapore and no open weights. Only companion repositories (Qwen-MM-Plugins, Qwen-Live-Harness) are open-sourced.
“Native omni model, API-only. Text, image, audio, and video in; text out on the standard Chat Completions / Responses API, with synthesized speech out on the realtime variant (WebSocket/WebRTC).”
— tokenstead.ai · source -
NEW Qwen3.8-Omni-Flash 🔥🔥 They're going FULL OMNI - Text, image, audio, and video inputs - 1M context "Audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash". Available only through API for now. This could be THE model
-
-
background
Alibaba releases Qwen 3.8 Omni Flash — Alibaba's Qwen team released Qwen 3.8 Omni Flash, a native omni model supporting text, image, audio, and video inputs with text and synthesized speech outputs via cloud API only.
Also covered reported alongside — the timeline has no entry for these yet
-
3 outlets Qwen 3.8 Omni Flash
first by HN Best, 8d ago · also HN Frontpage, The Decoder
2 more headlines
- Alibaba releases Qwen 3.8 Omni Flash HN Frontpage · 8d ago
- Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks The Decoder · 7d ago
What people are saying 17 voices from 1 site · best of 22 · verbatim
- Sep 18
-
Funny, that website ( https://qwen.ai/blog?id=qwen3.8-omni-flash ) downloads automatically 441 MiB of mp4 files... (using Firefox on Linux PC - noticed it because my network chart in Gkrellm spiked for several seconds).
-
The approach I normally take is: I have small benchmarks for myself which test for things I care about.And that has any models that I'm considering through that.
-
Pick the cheapest model with good speed. ZDR, and price. If it works great. If it doesn’t pick one a bit more expensive till you get what you need.
-
If you use the Chinese ones at least the energy comes from solar - aside from that, pick one and see if it solves your problems
-
>I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.Ask one of the top-tier models to do deep research on it.
-
For cost you can route the same task via openrouter and see what it costs. Literally a for m on models do task loop and look at the cost.
-
It's not like new releases come with fully mapped out capability scores for exactly the aspects that you're interested in. There are benchmarks, but reality is often different. It's simply unknown to humanity how well each model will perform in your own bespoke context unless you just try them. You can read experiences and vibes by others but…
-
I use claude for work, and opencode go at home, but i do dabble with openrouter from time to time, and then i just browse the model catalog and filter/sort by recent popularity, price, context size or whatever matters for the task.Mstly just popularity trends, hoping that there's some wisdom in the crowd.
-
The blog post itself is technically unimpressive, and feels like the slop future. Amongst many things - Scrolling in firefox is a nightmare (only uBlock origin) and the console is full or warnings and debug data. - Videos are all flashy but just fail to communicate anything beyond what can be said in a small paragraph (and with horrible stock…
-
Have you tried openrouters auto model? Tries to give you the best model based on prompt and price
-
My approach to this problem is to just...not try them all.As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe.So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't…
-
For TTS I launched this like this week based on a Reddit thread of recommendations, added a new one to it.. I want to say on Wednesday, but this week has been a blur. Problem I’ve found with similar sites is I can’t run a lot of the models, or the results are beyond stale.But this is only stuff I can run locally, or it’s a cloud model.So.. This is…
-
Please what is flash pro ultra and all these, can they just use semver or something
- Sep 17
-
If the performances are comparable, and there is no evidence it's not.in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47That is a massive cost reduction.Refs:
-
> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 FlashWow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.They also made a new harness but github link seems to 404.
-
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will…
-
Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.