Microsoft’s in-house model MAI-Image-2.5-Pro and MAI-Voice-2-Flash go live: GPU costs are cut significantly

Microsoft has released two new models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, and published internal data claiming that GPU costs can be cut by up to 89%. It also says its in-house models can replace OpenAI’s frontier models across multiple product lines, while packaging its training know-how as new Azure offerings for external sales.
(Background: Microsoft is rolling out 7 in-house models at once under the MAI umbrella, emphasizing compliance with “zero distillation,” and opening up model fine-tuning to attract enterprise users)
(Background add-on: Microsoft and OpenAI’s partnership is an “open marriage”—unusual, yet somewhat distant)

Table of contents

Toggle

  • The love triangle is built like this
  • Sell the training know-how as a product
  • Reactions split into two camps

On Wednesday, Microsoft disclosed a set of numbers: after the customer service center swapped its voice model to the in-house MAI-Voice-2-Flash, compute costs fell by 89%. At the same time, Microsoft also rolled out two new models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, entering public preview and targeting advanced image generation and large-scale voice processing, respectively.

The love triangle is built like this

This announcement of self-built models was actually foreshadowed long ago. Reuters reported in April this year that the original exclusive licensing arrangement between Microsoft and OpenAI has been rewritten into a non-exclusive setup; The Information also disclosed last September that Microsoft started importing Anthropic’s models into some products.

Microsoft CEO Nadella posted on X under the title “Frontier Diffusion and Control,” explaining the logic plainly: saturated frontier capabilities can now be delivered at lower cost via models optimized for high-frequency scenarios. “As long as our own models perform on par with or better than frontier replacement options, we start routing traffic to MAI.”

https://t.co/GUwbNABiw6

— Satya Nadella (@satyanadella) July 23, 2026

He also emphasized that “frontier models from OpenAI and Anthropic are still part of the coordinated system as a whole,” but at the same time proposed a model independence principle: in the evaluation system, “even if you remove any one model, it should still continue to climb.” In other words, keeping the supporting system, memory, and skills outside the model is precisely where Microsoft’s real control lies.

By this point, the love triangle has taken shape: Microsoft acts as the coordinator; OpenAI and Anthropic’s frontier models become interchangeable parts, while Microsoft’s own models capture an ever-growing share of everyday traffic.

Sell the training know-how as a product

What supports this narrative is another Microsoft announcement revealing the methodology: a “hill-climbing” strategy. Put simply, it’s about having the data, models, and “supporting system” around the product form a continuously iterating flywheel, rather than relying on a single training run to decide everything.

The most representative example is the lightweight model MAI-Code-1-Flash inside GitHub Copilot. Microsoft claims that in VS Code, its code adoption rate is about 10% higher than GPT-5.4 Mini and Claude Haiku 4.5, while using 10% fewer tokens.

Next, Microsoft “took it to the gym” by putting this model into the reinforcement learning environment in Excel. The result is a model small enough to run on older H100 GPUs, and even A100 GPUs, that in most Excel tasks it ties GPT-5.6—without needing the latest-generation chips—and it also allows the latest GB200 clusters to be redirected toward training rather than simply serving requests.

Microsoft didn’t keep this know-how in-house. Through Foundry and so-called “Frontier Tuning,” enterprise customers can train their own dedicated models using their own data, turning Microsoft’s cost-saving experiments into new Azure offerings directly. Microsoft also specifically stresses that these models are trained on “clean, traceable enterprise-grade data, without third-party model distillation.” In an industry environment where training data sources are increasingly scrutinized, this line is also meant for both enterprise customers and regulators.

Reactions split into two camps

After turning this know-how into a product, Microsoft also showed concrete scorecards. Bing Image Creator has fully switched to MAI-Image-2.5; on the PowerPoint side, GPU costs are down 84% compared with OpenAI’s GPT-Image-2. After OneDrive switches models, storage utilization rises by 26%, and latency drops by about 25%.

The most sensitive deployment is in healthcare: Dragon Copilot serves 170k healthcare professionals and handled 28 million patient records last quarter. After its cross-58-language transcription workflow switched to MAI-Transcribe-1.5, the transcription and language recognition error rate fell by 50% relative to before. As for the two new models themselves, MAI-Image-2.5-Pro is priced at $5 per million input tokens for text and $106 per million output tokens for images. In Microsoft’s announcement, advertising giant WPP’s global creative head Rob Reilly praised it as “a major leap forward for generative media tools.” MAI-Voice-2-Flash is twice as fast as its predecessor and 32% lower in cost, targeting scenarios like customer service centers—“high volume but not chasing top-end performance.”

But community reactions split into two camps. A designer on X criticized Microsoft for “historically doing the worst job when it comes to listening to user feedback,” arguing the company will lose the AI race as a result. Another user mocked Microsoft’s touted “model independence”: “After removing Microsoft, I’d actually like to see this system continue to climb.” Still, some developers supported it, saying, “Why does changing an Excel column require using omniscient large models?” And one line sums up the pitch: “Don’t dump everything into the biggest model—costs and performance will improve together.”

Skeptics’ doubts aren’t unfounded: adoption rates, storage rates, and GPU savings—these numbers all come from Microsoft’s own internal evaluations, not third-party benchmark tests, and Microsoft also controls which numbers it chooses to publish.

MSFT-2.24%
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned