A record published on August 27, 2026 describes an audio model combining speech recognition, audio understanding, speech synthesis, and editing.
On August 27, 2026, Radar recorded FireRedAudio, a model using Qwen3.5-9B as a shared backbone for automatic speech recognition (ASR), audio understanding, zero-shot and instruction-based speech synthesis, and speech editing. The description says it uses separate continuous representations for audio understanding and generation, and presents it as open source, covering multiple speech tasks in one audio language model.
To assess the result, consult the original publication and check the model documentation and materials to confirm supported tasks, requirements, and terms of use. The report provides no comparative metrics or performance details. If you use AI to study or apply this material, avoid sending identifiable organizational data until you have checked your organization's handling and access policies.