News
Microsoft releases high-fidelity image editing generation and low-latency text-to-speech model Microsoft today announced to further expand its self-developed AI model camp and officially launched the high-fidelity image generation and editing model MAI-Image-2
2 min read
Source: Telegram AI频道
Microsoft releases high-fidelity image editing generation and low-latency text-to-speech model Microsoft today announced to further expand its self-developed AI model camp, officially launching the high-fidelity image generation and editing model MAI-Image-2.5-Pro, and launching the public beta of the low-latency text-to-speech model MAI-Voice-2-Flash. As Microsoft's most accurate image generation model to date, MAI-Image-2.5-Pro is designed for application scenarios that require refined image generation, secondary editing, main visual poster design, and more accurate text rendering. In terms of pricing strategy, the model sells for US$5 per million text input Tokens, US$8 per million image input Tokens, and US$106 per million image output Tokens. Currently, developers and users can experience this model for free on the MAI Playground platform. At the same time, the MAI-Voice-2-Flash released by Microsoft focuses on high-concurrency, low-latency voice application scenarios such as customer service robots. Official data shows that while maintaining natural intonation and high-quality sound quality, MAI-Voice-2-Flash’s operating speed has doubled and its cost has been reduced by 32%. It is priced at US$15 per million characters. Microsoft revealed that many of its core businesses have been fully integrated into the self-developed MAI series models. For example, Bing Image Creator has fully switched to be driven by MAI-Image-2.5, and the graph-generating function of PowerPoint has also been connected to this model. Compared with the previously adopted third-party solution, the GPU cost has been reduced by up to 84% after switching to the self-developed model. In OneDrive, the application of MAI-Image-2.5 increases the image editing save rate by 26% and reduces P95 latency by approximately 25%. In addition, MAI-Voice-2-Flash has been first used in Dynamics 365 contact center, saving up to 89% of GPU costs, and is available as part of Azure Voice Live for developers to build voice interaction agents. The multilingual transcription model MAI-Transcribe-1.5 has also successfully provided Dragon Copilot with dictation services across 58 languages. via cnBeta.COM - Chinese industry information station (author: source: cnBeta.COM)