Microsoft Just Launched Seven AI Models at Once. One of Them Edits Images Better Than Almost Anything Else.
At Build 2026, Microsoft announced seven new in-house AI models covering reasoning, coding, transcription, voice, and image generation. But the one that matters most for anyone making visual content is MAI-Image-2.5. It just hit number two on the Arena image editing leaderboard, ahead of Nano Banana Pro, and number three for text-to-image generation. That is not a marginal improvement. That is a new top-tier option for anyone generating or editing images with AI.
The full model family spans text, code, voice, and images. MAI-Thinking-1 handles complex reasoning. MAI-Code-1-Flash is a coding assistant baked into GitHub Copilot. MAI-Transcribe-1.5 claims state-of-the-art accuracy across 43 languages. MAI-Voice-2 generates natural speech in 15 languages. All built from scratch by Microsoft, no distillation from third-party models, available on Azure Foundry, OpenRouter, Fireworks, and Baseten.
But MAI-Image-2.5 is the one that changes creative workflows. Here is what it actually does and how to use it.
What MAI-Image-2.5 Can Do
Two things make this model stand out: text-to-image quality and precision editing.
Text-to-image quality. MAI-Image-2.5 produces more detailed, coherent images from prompts. Text rendering inside images, the thing that has haunted every image model since DALL-E, is now genuinely usable. Microsoft reports a 107-point improvement on the text rendering Arena score compared to MAI-Image-2. Product imagery, packaging mockups, social graphics with overlaid text, all of this is now dramatically more reliable.
Precision editing. This is where MAI-Image-2.5 climbs to number two. You can tell it to change the colour of a bag in a photo, add flowers to a composition, or swap out text on a sign, and it edits only what you asked for. The rest of the image stays intact. No re-generation of the whole scene. No loss of composition. No starting over because the model decided to redraw the background.
Microsoft demonstrated this with a product shot where you could change a tote bag colour to orange, then add peonies peeking out of it, then adjust the text on the bag from one brand to another, and each edit preserved the original composition. That is the kind of control that previously required Photoshop skills or multiple regeneration attempts.
Face and identity consistency. The model preserves facial identity across edits. Change the pose, expression, or viewpoint and the person still looks like the same person. For anyone generating brand content with a consistent character or model, this is essential.
The Flash Variant: Fast and Cheap
MAI-Image-2.5-Flash is the speed and cost optimised version. It generates and edits at roughly half the price:
- MAI-Image-2.5: $5 per 1M text input tokens, $8 per 1M image input tokens, $47 per 1M image output tokens
- MAI-Image-2.5-Flash: $1.75 per 1M text input tokens, $1.75 per 1M image input tokens, $19.50 per 1M image output tokens
For bulk generation, product image variations, and batch creative testing, Flash is the one to use. For final output, presentations, and client-facing work, use the full model.
Where It Lives Inside Microsoft Products
MAI-Image-2.5 is already live in PowerPoint for generating presentation visuals. It is rolling out to OneDrive for photo editing: remove distractions, clean up backgrounds, enhance images, all while preserving the original scene.
If you use either of those apps, the model is already available. No API key needed. No separate subscription. It is embedded in the tools you already open.
How to Access It as a Developer
MAI-Image-2.5 is available on:
- Azure Foundry (Microsoft’s own platform)
- OpenRouter (immediate access for 9 million developers)
- Fireworks and Baseten (inference hosting)
You can also try it in the MAI Playground at microsoft.ai without writing any code.
For the first time, Microsoft is also allowing developers to tune the weights themselves. That means you can fine-tune MAI-Image-2.5 on your own product photography, brand style, or visual language and deploy it commercially.
What This Unlocks
- Product image editing without Photoshop. Change colours, add objects, swap text, remove backgrounds. All from a text prompt. No layer masks. No clone stamp. No starting over when the model changes something you did not ask for.
- Brand-consistent content at scale. Fine-tune the model on your brand assets, then generate variations that stay on brand. Every edit preserves identity. Every output follows your visual language.
- Presentation-ready visuals in PowerPoint. Generate slides from a prompt, then edit individual elements without regenerating the whole thing. The AI and the slide editor are the same interface.
- Photo cleanup in OneDrive. Remove unwanted objects, sharpen details, and adjust compositions without leaving your file storage. The editing comes to the photo, not the other way around.
- Bulk creative testing. Use the Flash variant to generate dozens of variations at a fraction of the cost. Test ad creatives, social graphics, and product shots at scale without burning through your budget.
What to Watch Out For
Like all image models, MAI-Image-2.5 can reflect biases in its training data and produce plausible but inaccurate visual details. Microsoft includes layered safety guardrails with prompt and output filtering, but generated images should be reviewed before use in identity, legal, medical, financial, or news contexts.
The editing precision is impressive but not perfect. Complex multi-element edits, like simultaneously changing text, swapping backgrounds, and adjusting lighting, may still require iteration. Start with single-change edits for the most reliable results.
Fine-tuning your own weights is powerful but requires infrastructure. If you are not already running custom models, start with the API first. Add fine-tuning when you have a clear use case that off-the-shelf generation does not cover.
The Bigger Picture: Seven Models, One Ecosystem
MAI-Image-2.5 is the headliner for creative professionals, but the full MAI family tells a bigger story. Microsoft built seven models from scratch, with zero distillation, across image, voice, transcription, coding, and reasoning. They share the same data pipeline and evaluation framework. They are designed to work together.
What that means in practice: you can generate an image with MAI-Image-2.5, build a presentation around it in PowerPoint, narrate it with MAI-Voice-2, transcribe the recording with MAI-Transcribe-1.5, and reason about the whole workflow with MAI-Thinking-1. All from one provider, one account, one billing system.
The most important shift, though, is Frontier Tuning. Microsoft is launching reinforcement learning environments that let you train MAI models on your own workflows. Your institutional knowledge becomes part of the model, and it stays yours. Early adopters are seeing 10x efficiency gains on tuned models versus general-purpose ones.
This is not just a model drop. It is a model ecosystem with a tuning path built in. The image model is the most immediately useful piece for anyone making visual content. But the infrastructure around it, the product integrations, the developer access, the fine-tuning, that is what makes it stick.
Want more like this? Join MOKU Club for free. Weekly resources, early access to new guides, and occasional templates you can actually use. Join below.



