On October 1, Microsoft AI (MAI) unveiled its first streaming transcription model, MAI-Transcribe-2-Streaming, along with the TTS model MAI-Voice-2.1 and its fast version, MAI-Voice-2.1-Flash. These models are available in public preview through Microsoft Foundry and are designed for building voice agents that interact with users.
MAI-Transcribe-2-Streaming offers low-latency transcription, returning interim results within just over 100 milliseconds and supporting automatic language detection for 60 languages, including Japanese. The pricing is set at $0.54 per hour of audio until the end of the year. Meanwhile, MAI-Voice-2.1 supports 23 languages and can switch languages while maintaining native accents, priced at $22 per million characters.
The two voice models can replicate voices from a few seconds of reference audio, requiring access approval and consent for use. The models are also available on platforms like MAI Playground and Vercel, with a demo voice agent named 'Chatter.' No further timeline was disclosed at the time of publication.
Editor's Note
Microsoft's introduction of the MAI-Transcribe-2-Streaming and TTS models reflects the growing competition in the voice AI sector, particularly against offerings from OpenAI and Google. As enterprises increasingly adopt voice technology for customer interaction, these advancements may influence procurement strategies and technology adoption in various industries.
Copyright Notice
This briefing is an independently written summary based on publicly available reporting and is provided for industry information and news discovery. The original report and source publication are credited and linked where applicable. RobotToday does not claim ownership of third-party source material.
Rights concerns? If you believe any material in this briefing infringes your copyright or other rights, please contact [email protected] with the relevant URL and details. We will review the matter and take appropriate action where warranted.
Leave a comment