Microsoft introduces the MAI-Transcribe-2 speech model
Announced on September 3, 2026, MAI-Transcribe-2 is a transcription model reporting 2.0% AA-WER, about 411× real-time speed, and support for 60 languages.
On September 3, 2026, Microsoft announced MAI-Transcribe-2, a speech recognition model that the announcement says records 2.0% AA-WER and processes audio at about 411 times real time. It supports 60 languages, speaker diarization, word-level timestamps, and clean or verbatim transcription styles.
The announcement suggests that engineers compare reported accuracy, speed, and price when evaluating speech recognition services. Consult Microsoft's original announcement to verify the specifications and their conditions; results and costs may depend on the use case. If using AI to analyze evaluation materials, avoid entering personal data or confidential content unless your organization's protection policy is applied.