News

Microsoft releases high-fidelity image editing generation and low-latency text-to-speech model Microsoft today announced to further expand its self-developed AI model camp and officially launched the high-fidelity image generation and editing model MAI-Image-2

2 min read
Microsoft releases high-fidelity image editing generation and low-latency text-to-speech model Microsoft today announced to further expand its self-developed AI model camp, officially launching the high-fidelity image generation and editing model MAI-Image-2.5-Pro, and launching the public beta of the low-latency text-to-speech model MAI-Voice-2-Flash. As Microsoft's most accurate image generation model to date, MAI-Image-2.5-Pro ​​is designed for application scenarios that require refined image generation, secondary editing, main visual poster design, and more accurate text rendering. In terms of pricing strategy, the model sells for US$5 per million text input Tokens, US$8 per million image input Tokens, and US$106 per million image output Tokens. Currently, developers and users can experience this model for free on the MAI Playground platform. At the same time, the MAI-Voice-2-Flash released by Microsoft focuses on high-concurrency, low-latency voice application scenarios such as customer service robots. Official data shows that while maintaining natural intonation and high-quality sound quality, MAI-Voice-2-Flash’s operating speed has doubled and its cost has been reduced by 32%. It is priced at US$15 per million characters. Microsoft revealed that many of its core businesses have been fully integrated into the self-developed MAI series models. For example, Bing Image Creator has fully switched to be driven by MAI-Image-2.5, and the graph-generating function of PowerPoint has also been connected to this model. Compared with the previously adopted third-party solution, the GPU cost has been reduced by up to 84% after switching to the self-developed model. In OneDrive, the application of MAI-Image-2.5 increases the image editing save rate by 26% and reduces P95 latency by approximately 25%. In addition, MAI-Voice-2-Flash has been first used in Dynamics 365 contact center, saving up to 89% of GPU costs, and is available as part of Azure Voice Live for developers to build voice interaction agents. The multilingual transcription model MAI-Transcribe-1.5 has also successfully provided Dragon Copilot with dictation services across 58 languages. via cnBeta.COM - Chinese industry information station (author: source: cnBeta.COM)