Qwen3-Omni combines speech, text, images and video
On September 22, 2025, Qwen announced the release of three multimodal models that can follow instructions, reason and generate captions, with speech input and output.
On September 22, 2025, Qwen said it had released three Qwen3-Omni models. According to the announcement, they follow instructions, reason and generate captions; combine text, images, audio and video; and accept and produce speech.
The announcement highlights their relevance to engineers interested in exploring an open-source model that processes speech alongside other modalities. To verify the scope and technical details, consult Qwen’s original announcement and compare its claims with the model documentation and materials; the available text gives no benchmark results.