Skip to content
Rota Nacional

Radar ·

Step-Audio-EditX: prompt-based audio editing

Announced on November 8, 2025, the model is described as a 3-billion-parameter system for audio editing, zero-shot multilingual speech synthesis, and prompt-based vocal control.

On November 8, 2025, StepFun announced Step-Audio-EditX, described as a 3-billion-parameter model for audio editing. According to the post, it also offers zero-shot multilingual speech synthesis and prompt-based control of emotion, speaking style, and vocal elements such as breaths and laughter. The announcement states that it can run on a single GPU and is licensed under Apache 2.0.

The proposal is to evaluate one model for speech generation and iterative audio editing. These capabilities and conditions are claims in the announcement; to verify them, consult the original post and project materials, check the license, and test results with your own examples. If using AI to study or apply the approach, avoid submitting identifiable recordings without authorization and review your organization's data-handling rules.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free