Published on January 12, 2026, the announcement describes a multilingual audio dataset with 150,000 hours, coverage of 10 languages and word-level timestamps, along with the generative models LEMAS-TTS and LEMAS-Edit.
On January 12, 2026, a LEMAS publication announced a multilingual speech dataset with 150,000 hours of audio. According to the announcement, it covers 10 languages, includes word-level timestamps, and is accompanied by the generative models LEMAS-TTS and LEMAS-Edit.
The material points to possible uses in multilingual text-to-speech (TTS) and speech editing. To check the scope and details, consult the original LEMAS publication and verify its information about the data and models; the summarized announcement does not specify the languages or provide evaluation results.