Microsoft has introduced three artificial intelligence models for real-time speech recognition and voice generation, expanding its push into conversational AI as businesses increasingly explore automated customer service and contact centre applications.
The models, MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, are designed for developers building voice agents, live transcription services, customer support systems and other applications that require AI to process spoken conversations with minimal delay.
MAI-Transcribe-2-Streaming converts speech into text while a person is still speaking rather than waiting for an entire sentence or recording to finish. It supports continuous transcription across 60 languages and can automatically identify language changes during conversations.
Microsoft said the model can generate its first transcription hypotheses within the low hundreds of milliseconds. Once an utterance is complete, it produces a stable transcript that applications can use for subsequent actions.
The capability could be particularly relevant to contact centres, where reducing the delay between a customer's question and an automated response is important. A voice agent could begin interpreting an enquiry, preparing an answer or triggering another system while the conversation is still underway.
Alongside transcription, Microsoft has introduced MAI-Voice-2.1 for generating speech. The multilingual text-to-speech model supports 23 languages and is designed for applications where more natural and expressive AI-generated voices are required.
MAI-Voice-2.1-Flash supports the same language coverage and cross-language voice identities but prioritises speed, scale and cost. This gives developers a choice between a model focused on expressive speech and a faster alternative intended for applications handling larger volumes of interactions.
Microsoft said MAI-Transcribe-2-Streaming ranked first for both final and partial transcript accuracy in evaluations published by benchmarking platform Artificial Analysis. As with other benchmark results, actual performance can vary depending on accents, audio quality, language and deployment conditions.
The models add another layer to Microsoft's existing customer service and contact centre technology. Copilot Studio already allows organisations to develop real-time voice agents, while Dynamics 365 Contact Center combines AI with customer service workflows.
Microsoft has also been adding capabilities such as real-time translation, voice biometrics and customised neural voices to its enterprise communications portfolio.
The latest release gives developers additional speech infrastructure for building AI agents that can listen and respond during live conversations. For businesses, the models could extend the use of generative AI beyond text-based chatbots into voice interactions across customer service, sales and other conversational applications.
Disclaimer: This article may include information derived from interviews, press releases, public statements, research, company communications and other publicly available or third-party sources. Such material may be summarised, paraphrased or contextualised for journalistic and editorial purposes. All rights in third-party content remain with their respective owners.