News
Microsoft self-developed dual model MAI-Image-2
3 min read
Source: Telegram AI频道
Microsoft's self-developed dual models MAI-Image-2.5-Pro and MAI-Voice-2-Flash are released: without distilling third parties, they have been implemented in Bing and PowerPoint. On July 23, Microsoft announced the launch of two self-developed AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, both of which have entered the public preview stage. The two models are respectively aimed at high-quality image generation and high-concurrency voice interaction scenarios. Microsoft emphasizes that its training data has been cleaned and traceable, does not rely on third-party model distillation, and is designed to directly serve Microsoft product users from the bottom up. The image model has the highest accuracy and reduces GPU costs by up to 84%. MAI-Image-2.5-Pro is Microsoft's most accurate image model at present. It is designed for high-demand scenarios such as main visual image generation, detail editing, and text rendering in images, and supports natural language editing instructions. This model has become the default image generation model in Bing Image Creator and is used for image-to-image editing functions in PowerPoint, reducing GPU costs by up to 84% compared to GPT-Image-2. After deployment in OneDrive, the success rate of saving related scenes increased by 26%, P95 latency was reduced by approximately 25%, and the efficiency of medium-load production environments was increased by 2.5 times. In terms of pricing, text input costs $5 per million Tokens, and image output costs $106 per million Tokens. The voice model speed is increased by 2 times, and it has served customers such as T-Mobile. MAI-Voice-2-Flash is optimized for high-frequency voice applications. Compared with the previous generation, the speed is increased by about 2 times and the cost is reduced by 32%. It is suitable for fast response scenarios such as customer service centers and voice assistants. This model has been connected to Dynamics 365 Contact Center to provide services to customers such as T-Mobile and EasyJet, with GPU costs reduced by up to 89%. Microsoft has also integrated it into the Azure Voice Live service to support developers in building voice-to-voice interactive agents. In addition, MAI-Transcribe-1. 5 of the same series has been applied to the medical voice solution Dragon Copilot, which is used by 170,000 medical service providers. It processed 28 million patient consultation records last quarter and supports 58 languages. The transcription error rate in most languages has been reduced by 50%. Microsoft revealed that the next generation GB200 computing cluster has been put into operation, and the MAI series models will continue to expand. via AI News (author: