Skip to content
Rota Nacional

Radar ·

MaskGCT: zero-shot TTS with masked codec

A paper published on October 24, 2024 presents MaskGCT, a zero-shot text-to-speech model. The publication says it can clone a voice from five seconds of speech.

Published on October 24, 2024, the paper presents MaskGCT, a zero-shot text-to-speech (TTS) model based on a masked codec Transformer. According to the publication, the model can clone a voice using five seconds of speech. The project was made available by Amphion; this result is the publication’s claim, not an independent verification in this summary.

Engineers interested in speech synthesis can consult the paper and project materials to examine the approach and voice-cloning requirements. Check the original materials for conditions, evaluation methods, and limitations before drawing conclusions or testing the system. If using AI to study or apply the technique, do not submit identifiable recordings without authorization, and review your organization’s voice-data policies.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free