Microsoft Adds Fast MAI-Image Model to Foundry
MAI-Image-2.6-Flash pairs lower prices with production-focused editing, while the full model leads a major third-party benchmark.
Two models for production image workloads
Microsoft has introduced MAI-Image-2.6 and the faster MAI-Image-2.6-Flash in public preview through Microsoft Foundry. The release positions the company’s in-house image family as a production alternative for businesses that need image generation and editing at predictable cost and speed.
MAI-Image-2.6 adds multi-reference editing, web grounding and expanded control over output format and resolution. Microsoft says it improves text rendering, instruction following and visual consistency over MAI-Image-2.5. The Flash variant carries those capabilities into a lower-cost model intended for higher-volume applications; Microsoft reports that it is more than twice as fast as GPT-Image-2-Medium and 78% more efficient, although those vendor comparisons still need independent workload testing.
Foundry pricing starts at $38 per million image-output tokens for MAI-Image-2.6 and $19 for Flash, with separate charges for text and image inputs. On Artificial Analysis’s image-editing leaderboard, the full model ranked first with an Elo score of 1,325, while Flash scored 1,311 and occupied a statistically overlapping group near the top. The leaderboard listed estimated costs of approximately $38.90 and $19.50 per thousand images respectively, compared with about $211 for GPT Image 2 at its high setting.
Why it matters
Microsoft is competing on the part of image generation that enterprise buyers can measure: editing consistency, throughput and unit economics. The Flash model is especially consequential because it narrows the quality gap while roughly halving the full model’s output price. Public-preview status means reliability, regional availability and behavior on brand-sensitive workflows remain unsettled, but Microsoft now has an image stack it can bundle directly with Foundry governance and enterprise procurement rather than relying entirely on an external model provider.