RELEASE · MODELS · #679
Qwen launches Qwen3.8-Omni-Flash multimodal agent model and undercuts Gemini Flash pricing
Qwen announced Qwen3.8-Omni-Flash, its first multimodal agent-focused model with a one‑million‑token context window that processes audio and video together, performs tasks like vlog editing and movie summarization, and supports real‑time interaction via Qwen‑Live Harness. API pricing is $0.15 per million input tokens and $0.47 per million output tokens (Qwen estimates audio input under $0.01/hour and 720p video at 1 fps about $0.20/hour), and the model is available through Qwen Studio, Qwen Cloud and the Qwen API; Qwen also published open-source Qwen-MM-Plugins for agent integrations.
KEY POINTS
- Qwen announced Qwen3.8-Omni-Flash, its first multimodal agent-focused model with a one‑million‑token context window that processes audio and video together, performs tasks like vlog editing and movie summarization, and supports real‑time interaction via Qwen‑Live Harness.
- API pricing is $0.15 per million input tokens and $0.47 per million output tokens (Qwen estimates audio input under $0.01/hour and 720p video at 1 fps about $0.20/hour), and the model is available through Qwen Studio, Qwen Cloud and the Qwen API; Qwen also published open-source Qwen-MM-Plugins for agent integrations.
- Matching Gemini 3.8 Flash on multimodal benchmarks while offering far lower API pricing and a 1M‑token context window could make advanced multimodal agent capabilities more accessible and cost‑effective for developers.
WHY IT MATTERS
Matching Gemini 3.8 Flash on multimodal benchmarks while offering far lower API pricing and a 1M‑token context window could make advanced multimodal agent capabilities more accessible and cost‑effective for developers.