Google's Gemini 3.8 Flash charges $0.75 per million input tokens and $3.75 per million output tokens, and that’s the introductory rate. Qwen just put a multimodal model on the market at $0.15 and $0.47.
Qwen3.8-Omni-Flash is Qwen’s first multimodal model built for AI agents, and on audio-video tasks the company says it comes close to matching Gemini 3.8 Flash. The pitch is close-enough performance at a fraction of the bill.
What the pricing looks like on a real job
Qwen puts audio input at under $0.01 per hour. Feed it 720p video with audio at one frame per second and the estimate is about $0.20, though that figure doesn’t include what you pay for the response.
Gemini’s numbers move the other way. Google’s introductory rate is set to double on January 1, 2027, so the gap Qwen is advertising today gets wider if Google follows through.
Built to run agents, not just answer prompts
The model processes audio and video together, draws conclusions from them and uses tools on its own. Qwen’s examples are editing vlogs, translating short videos and summarizing movies. The context window spans one million tokens.
I haven’t run it against a production workload yet, so the benchmark claim is Qwen’s, not mine. And the company’s own framing is that it comes close to matching Gemini 3.8 Flash, which is a hedge worth reading twice before you migrate anything.
Where you can get it
Qwen3.8-Omni-Flash is available through Qwen Studio, Qwen Cloud and the API. The open-source Qwen-MM-Plugins bolt video editing, speaker recognition, PDF video notes and reusable workflows onto agents including Claude Code, Gemini CLI and Qwen Code. Qwen-Live Harness handles real-time interaction through a camera and microphone.
If you’re already paying Gemini 3.8 Flash’s output rate for video summaries, the $0.47 figure is the one to check first.