
MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is suited for voice agents, assistants, call centers, and other interactive applications where latency and cost matter most.
On OpenRouter, set voice to a full voice ID with the model suffix, such as "en-US-Harper:MAI-Voice-2.1-Flash". A voice's locale sets the synthesis language. Set response_format to "mp3" or "pcm" (24 kHz mono). Harper and Grant support the agent, customer-call-center, educational, and narrator speaking styles, and many locale voices add emotion styles such as excited, happy, sad, and whispering. The full list of voices is in the supported_voices field of the models APIOpens in new tab. See the text-to-speech guideOpens in new tab.
| $15.00 | 0.61s |
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.