CocktailASR-1 transcribes a selected speaker in mixed audio
Xiaomi's model uses a reference voice to transcribe a selected speaker in overlapping audio; the publication reports support for English and Chinese.
CocktailASR-1 is described as a speaker-directed speech recognition model. Given a reference voice, it aims to transcribe the selected voice even when speech overlaps. The publication reports support for English and Chinese.
The result may interest engineers evaluating speaker-conditioned speech recognition. If you use AI to study or apply the technique, protect recordings and transcripts: obtain authorization, restrict access, and avoid sending personal data unnecessarily. Audio transcription at Rota passes through the personal-data barrier and organizational policies; this does not mean the platform implements this model or solves the engineering problem.