Microsoft Releases New AI Models for Image Generation and Voice Processing
Microsoft announced the launch of two new proprietary AI models, MAI-Image-2.5-Pro for image generation and MAI-Voice-2-Flash for voice processing. Both models are now in public preview and have been integrated into several Microsoft services.

Microsoft unveiled two new proprietary artificial intelligence models on July 23: MAI-Image-2.5-Pro and MAI-Voice-2-Flash. Both models are now available in public preview and are designed for high-quality image generation and efficient voice interaction, respectively.[br]The company emphasized that the new models have been trained using clean, traceable, enterprise-grade data without relying on third-party model distillation. MAI-Image-2.5-Pro is described as Microsoft's most accurate image model to date, suitable for demanding applications such as generating key visuals, detailed editing, and rendering text within images, controllable through natural language prompts.[br]MAI-Voice-2-Flash is optimized for high-concurrency voice applications. It is reported to be approximately twice as fast and 32% cheaper than its predecessor, while maintaining natural speech tone and high-quality output with low latency.[br]Both models are already being integrated into Microsoft's product ecosystem. Bing Image Creator now uses MAI-Image-2.5-Pro as its default image generation model. In PowerPoint, the model powers image editing features, and OneDrive utilizes it for certain image manipulation functions. MAI-Voice-2-Flash has been deployed in Dynamics 365 Contact Center for enterprise customer service AI agents and integrated into Azure Voice Live for developers.[br]Microsoft stated that the MAI series of models will continue to expand and be offered to developers through the Foundry platform. The company has also commissioned its next-generation GB200 computing clusters to support ongoing model development.