# Microsoft introduces MAI-Voice-2 with multilingual voice controls

> MAI-Voice-2 adds multilingual speech generation, emotion tags and reference-audio prompting, creating new evaluation work for voice applications.

Canonical URL: https://www.devobs.io/news/news-mai-voice-2-launch/
By: Lucas Vale
Published: 2026-09-06T11:58:54.629Z
Updated: 2026-09-06T11:58:54.629Z
Event date: 2026-06-02
Section: AI

Microsoft introduced MAI-Voice-2 on June 2, expanding its text-to-speech model with multilingual generation, emotion controls and reference-audio prompting. The [announcement](https://microsoft.ai/news/mai-voice-2/) said the model was available in Microsoft Foundry.

## More controls over generated speech

Microsoft described emotion tags, support for selected mixed-language combinations and an emphasis on maintaining a speaker’s identity through longer recordings. The release also described consent guardrails around voice prompting. These are the company’s product statements; they do not establish that every deployment automatically meets an organization’s consent or quality requirements.

For teams building spoken interfaces, the change introduces more settings that can affect how a message is perceived.

## Review meaning as well as sound

Our recommended test set would include questions, warnings, numbers, names and deliberately neutral information. Reviewers should assess whether the delivery matches the text’s purpose. An enthusiastic tone may fit a tutorial but be inappropriate for a billing correction or an account-recovery instruction.

Test complete passages as well as isolated sentences. A sample that sounds convincing for ten seconds does not answer whether terminology, pacing and speaker character remain acceptable across a longer lesson or narrated document. Mixed-language material should receive review from listeners who understand the actual combination being used.

## Treat voice assets as governed inputs

Reference recordings should have a documented owner, approved purpose and retention decision before they become reusable production inputs. Keep those decisions attached to the voice configuration so a later editor does not have to reconstruct permission from a loose audio file.

Finally, retain an accessible text version and an obvious way to stop or replay speech. Those interface choices remain the application team’s responsibility even when a model supplies more natural and controllable audio.

## Source references

- <https://microsoft.ai/news/mai-voice-2/>
