Microsoft announces long-audio transcription with VibeVoice-ASR
On March 2, 2026, Microsoft said VibeVoice-ASR transcribes audio up to 60 minutes in one pass and offers speaker diarization, timestamps, hotwords, and support for more than 50 languages.
On March 2, 2026, Microsoft announced that VibeVoice-ASR can transcribe audio up to 60 minutes in one pass. According to the company, the system also identifies speakers, associates timestamps with the content, accepts hotwords, and supports more than 50 languages without requiring the language to be set in advance.
The announcement points to possible uses in multilingual speech-recognition pipelines, but this summary does not provide independent evaluation results. To check the claims, consult Microsoft’s original announcement and test the system with representative audio, checking languages, timestamps, and speaker identification. If you use AI to analyze recordings, protect personal data according to your organization’s policy.