NewsStocksMicrosoft Previews MAI-Realtime, a Bidirectional Voice Model for Conversational AI

Microsoft Previews MAI-Realtime, a Bidirectional Voice Model for Conversational AI

Author: CryptoBriefing·

Key Takeaways

  • MAI-Realtime is a bidirectional voice model that processes audio in both directions simultaneously, enabling full-duplex conversations rather than the rigid turn-taking of traditional voice systems.
  • Microsoft has assembled a vertically integrated voice AI stack consisting of MAI-Transcribe-1 for speech recognition, MAI-Voice-1 for text-to-speech, and MAI-Realtime for real-time bidirectional conversation.
  • MAI-Voice-1 can generate a full minute of audio in under one second on a single GPU and has already been integrated into Copilot features across multiple Microsoft applications.
  • By building its own voice models, Microsoft avoids paying third-party inference costs and gains full control over its product roadmap and pricing decisions.
  • Microsoft has not yet opened MAI-Realtime to public testing, released performance benchmarks, or confirmed specific plans for integrating it into Copilot or other products.
Microsoft Previews MAI-Realtime, a Bidirectional Voice Model for Conversational AI

Microsoft is steadily constructing its own voice AI stack, and the newest component is MAI-Realtime, a bidirectional voice model currently in internal preview. The company states that the model will deliver more natural interactions across multiple languages and integrate with its broader tool ecosystem.

Unlike traditional voice systems that alternate between speaking and listening in a walkie-talkie fashion, a bidirectional model handles both simultaneously. MAI-Realtime processes audio in real time in both directions, enabling interactions that feel conversational rather than rigid or robotic. This full-duplex approach mirrors how human conversation actually works — people interrupt, overlap, and respond mid-sentence — and has been a technical frontier for voice AI, where latency and turn-taking logic have long made interactions feel mechanical.

MAI-Realtime remains in an internal preview phase, meaning Microsoft has not yet opened it to public testing or published performance benchmarks. What the company has confirmed is that the model is designed to feel more natural, support multiple languages, and work alongside existing tools — which, within Microsoft's ecosystem, likely points to integration with Copilot and its suite of productivity applications.

MAI-Realtime does not exist in isolation. Microsoft has been assembling a full voice AI stack under its MAI (Microsoft AI) umbrella. MAI-Voice-1, the company's expressive text-to-speech model, was first previewed in August 2025 and expanded in April 2026. That model can produce a full minute of audio in under one second using a single GPU and has already been integrated into Copilot features across multiple Microsoft applications. Alongside it, Microsoft introduced MAI-Transcribe-1, a speech recognition model supporting 25 languages.

Taken together, the three models form a vertically integrated voice AI stack: MAI-Transcribe-1 handles listening, MAI-Voice-1 handles speaking, and MAI-Realtime enables full bidirectional conversation — all owned by Microsoft end to end. For Microsoft, which counts hundreds of millions of enterprise users across Teams, Office, and Windows, controlling this stack internally means it can embed voice-based AI interactions deeply into workflows — from live meeting transcription and summarization in Teams to voice-driven document drafting — without depending on a third-party provider's release cadence or pricing.

This proprietary approach carries clear strategic implications. When Microsoft relies on OpenAI's voice capabilities, it pays for inference, remains subject to OpenAI's pricing decisions, and does not control the product roadmap. Building its own models changes that calculus entirely.

The competitive landscape is active. OpenAI released its own real-time voice API in late 2024, and Google has been advancing Gemini's multimodal capabilities, which include audio processing. Amazon has also invested heavily in voice AI through Alexa and its Bedrock platform. The common thread across these efforts is a push toward voice as a primary interface for AI — not just a feature, but a fundamental way users interact with models. Microsoft's move signals that the company intends to compete on this axis with its own technology rather than licensing a partner's.

If MAI-Realtime is eventually folded into Copilot for Microsoft 365, it would reach enterprise users across Word, Excel, Teams, and Outlook. Real-time voice interaction in these tools could lower the barrier for users who find typed prompts cumbersome, though Microsoft has not confirmed specific product plans.

Key open questions include whether MAI-Realtime progresses from internal preview to public developer access, how quickly it integrates into Copilot's existing voice features, and whether Microsoft begins publishing benchmark comparisons against OpenAI's real-time voice API.