The model is designed to handle everything from meeting transcripts and captions to legal documentation and other real-world audio applications.
Model Supports 60 Languages
According to Microsoft, MAI-Transcribe-2 ranks first on the FLEURS benchmark across 60 languages, with an average word-error rate of 5.2%. The company also claims the model can process audio up to 10 times faster than some competing systems.
The model includes speaker identification, word-level timestamps, automatic language detection and keyword recognition for specialized terminology. It can also handle code-switching, including conversations that mix languages such as English and Hindi.
Microsoft Targets Lower Costs
Microsoft is also highlighting the model’s pricing. MAI-Transcribe-2 will launch at $0.10 per hour of audio as a limited-time offer through the end of 2026.
The model is available through Microsoft Foundry, MAI Playground and OpenRouter, giving developers access to the technology for applications ranging from accessibility and captioning to enterprise transcription and voice-based services.

Reader Discussion
Join 0 thoughts shared by the communityBe the First to Comment
No discussions started yet. Share your feedback or insights with the TechCrest community!