Skip to content
Rota Nacional

Radar ·

LEMAS announces speech dataset with 150,000 hours

Published on January 12, 2026, the announcement describes a multilingual audio dataset with 150,000 hours, coverage of 10 languages and word-level timestamps, along with the generative models LEMAS-TTS and LEMAS-Edit.

On January 12, 2026, a LEMAS publication announced a multilingual speech dataset with 150,000 hours of audio. According to the announcement, it covers 10 languages, includes word-level timestamps, and is accompanied by the generative models LEMAS-TTS and LEMAS-Edit.

The material points to possible uses in multilingual text-to-speech (TTS) and speech editing. To check the scope and details, consult the original LEMAS publication and verify its information about the data and models; the summarized announcement does not specify the languages or provide evaluation results.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free