Microsoft previews MAI image and voice models in Foundry

Editorial graphic for Microsoft's MAI image and voice model previews.LinkLoot editorial graphic
Editorial graphic for Microsoft's MAI image and voice model previews.LinkLoot editorial graphic
AI & Automation

Microsoft AI has put MAI-Image-2.5-Pro and MAI-Voice-2-Flash into public preview, giving Foundry builders new first-party options for high-fidelity image work and lower-cost voice generation.

Microsoft AI has moved two more in-house media models into public preview: MAI-Image-2.5-Pro for higher-fidelity image generation and editing, and MAI-Voice-2-Flash for faster, cheaper speech generation. Both are positioned for Microsoft Foundry builders rather than as standalone consumer toys.

The release matters because Microsoft is filling more of its own AI stack with first-party models. The announcement says MAI models already power production surfaces across Bing, PowerPoint, OneDrive, Dynamics 365, and Azure, while the new previews give developers direct access to the quality-speed-cost tradeoffs behind those products.

MAI-Image-2.5-Pro targets high-fidelity creative work

MAI-Image-2.5-Pro is the quality-focused branch of Microsoft’s image family. Microsoft describes it as its highest-fidelity image model so far, aimed at hero imagery, detailed editing, and more accurate in-image text rendering.

The Foundry model catalog lists the model as a preview option for text-to-image and image-to-image workflows. Inputs can include text and JPEG or PNG images for editing. Output is a single PNG image, with a 32,000-token context length listed in Microsoft Learn. Image dimensions start at 768 by 768 pixels, with a maximum total pixel count of 1,048,576.

Pricing in Microsoft’s announcement is explicit: $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens. That makes this a model teams should route deliberately, especially when the job needs precise visual structure or text handling rather than high-volume draft generation.

MAI-Voice-2-Flash is built for scaled voice agents

MAI-Voice-2-Flash takes the opposite side of the tradeoff. Microsoft says it keeps the natural prosody and acoustic quality of MAI-Voice-2 while running 2x faster and 32% cheaper. The listed price is $15 per 1M characters.

The practical audience is high-volume voice work: contact centers, voice agents, and applications where latency and unit cost matter more than squeezing out the most expressive generation every time. Microsoft says the model powers Dynamics 365 Contact Center and is integrated into Azure Voice Live, which gives developers a speech-to-speech path for voice agents.

That product placement is the strongest signal. This is not only a model-card update; Microsoft is using the same family inside production products, then exposing preview variants through Foundry so enterprises can build against them.

Foundry access comes with preview caveats

Foundry availability should still be treated as a preview. Microsoft Learn marks MAI-Image-2.5-Pro as Preview, and model availability can vary by region, deployment type, and subscription quota. Teams should confirm the exact model ID, supported endpoint, region, and billing behavior inside their own Azure tenant before committing workflows.

The model split also points to a routing strategy. Use MAI-Image-2.5-Pro when visual accuracy, in-image text, or surgical editing is worth the premium. Use faster or cheaper image models for drafts, thumbnails, and internal exploration. Use MAI-Voice-2-Flash when responsiveness and scale drive the business case.

For teams comparing media models, keep a small internal benchmark: one branded image edit, one prompt with embedded text, one long voice generation, and one low-latency agent response. That will show whether Microsoft’s preview models fit your actual workload better than generic leaderboard claims.

Sources and methodology

The primary source is Microsoft AI’s July 23 announcement for MAI-Image-2.5-Pro and MAI-Voice-2-Flash. LinkLoot cross-checked Microsoft Learn’s Foundry model catalog for model capabilities and used Neowin as independent publication context. The claims above preserve Microsoft’s preview wording and pricing qualifiers rather than treating the models as generally available.