MiniMax releases H3 for unified 2K video and native stereo audio
MiniMax H3 combines text, images, video, and audio in one generation context, with 2K output, native stereo sound, reference-based creation, and natural-language video editing.
MiniMax has released H3, a general-purpose multimodal video model that accepts text, images, video, and audio in one context. The model generates clips in 2K with native stereo sound and supports text-to-video, first-and-last-frame control, multimodal reference generation, and editing of existing footage through natural-language instructions.
This is a broader creative system than a conventional text-to-video endpoint. MiniMax positions the same model for generation, reference transfer, and targeted edits, while launch partner fal.ai is already exposing hosted H3 endpoints for developers.
MiniMax H3 puts four input types into one context
H3 can combine multiple reference types inside one request. MiniMax's official documentation allows up to nine images, three video clips, and three audio clips, with a maximum of 12 files. A prompt can contain up to 7,000 characters.
That input design gives creators several control paths in one generation. An image can establish a character or product, a video can provide motion or camera behavior, an audio clip can provide a voice or sound reference, and text can define the requested scene and edits. The model is designed to preserve referenced subjects and assets across the generated clip.
| Capability | MiniMax H3 launch specification |
|---|---|
| Output | 2K video with native stereo audio |
| Duration | 4 to 15 seconds in MiniMax's API documentation |
| Inputs | Text, images, video, and audio |
| Reference limits | Up to 9 images, 3 videos, and 3 audio clips; 12 files total |
| Prompt limit | Up to 7,000 characters |
| Modes | Text-to-video, first/last frame, reference generation, video editing |
Native audio and targeted editing expand the production scope
Every H3 generation includes stereo audio rather than requiring a separate sound-generation pass. MiniMax and fal.ai describe output that can include dialogue, score, ambience, and effects synchronized to the picture. Reference audio can also guide voice transfer, subject to the rights and consent required for the supplied material.
The editing workflow is equally significant. H3 is designed to replace or remove objects, change backgrounds or lighting, rewrite visible text, alter dialogue, and modify pacing while keeping untouched regions stable. Those controls target practical commercial work such as product films, social clips, game visuals, interface motion, branded typography, and localized campaign variants.
Hosted access is live, while open-weight details need verification
MiniMax documents MiniMax-H3 on its video-generation API, which uses an asynchronous task flow: submit a request, poll the task ID, and retrieve the resulting video URL. fal.ai also offers serverless text-to-video, image-to-video, and reference-to-video endpoints, so developers can evaluate H3 without provisioning their own inference hardware.
fal.ai labels H3 as an open-weights model, and MiniMax's announcement says model weights are part of the release direction. Teams planning self-hosting should still verify the actual weight artifact, license, hardware requirements, and permitted commercial use before committing infrastructure. Hosted API access and downloadable, production-ready weights are separate availability milestones.
What creators and developers should test first
The headline specifications do not answer every production question. A useful evaluation should focus on repeatability: whether identities and products remain consistent across shots, whether localized edits leave the rest of a scene intact, and whether generated dialogue and effects stay synchronized over the full clip.
- Test the same subject across text-only, image-reference, and mixed-reference requests.
- Compare first-and-last-frame control with full reference generation for transitions.
- Stress-test typography, logos, interfaces, and product geometry rather than judging only cinematic samples.
- Review consent, likeness, voice, trademark, and music rights before using references commercially.
- Measure queue time and total cost at the exact duration, resolution, and endpoint required by the workflow.
MiniMax H3 is available now through MiniMax's documented API path and fal.ai's hosted endpoints. The next important milestone is a clearly verifiable weight release with complete licensing and deployment guidance; until then, the hosted services are the most concrete route for testing its unified audiovisual workflow.
