Skip to content
Rota Nacional

Radar ·

VoiceCraft: speech editing and zero-shot TTS

Published on August 25, 2024, VoiceCraft is described as a neural codec language model for speech editing and zero-shot text-to-speech. The publication says it can clone or edit an unseen voice using a few seconds of reference audio.

Published on August 25, 2024, VoiceCraft is presented as a neural codec language model for speech editing and zero-shot text-to-speech (TTS). According to the publication, it can clone or edit an unseen voice using a few seconds of reference audio.

The text points to the approach as something engineers can evaluate on varied audio, but provides no test results or performance measurements. To verify the scope and evidence, consult the original publication and examine the materials and evaluations it provides.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free