
Table of Content

Try DomoAI, the Best AI Animation Generator
Turn any text, image, or video into anime, realistic, or artistic videos. Over 30 unique styles available.
MiniMax H3 is MiniMax's general-purpose, omni-modal system for generating video with native stereo audio from text, image, video, and audio context. MiniMax M3 is a separate coding-and-agent model. Hailuo AI is a user-facing product surface, while Hailuo 2.3 is an earlier video model.
H3 is a new release, so access details and implementation-specific controls may change.
MiniMax launched H3 on July 31, 2026. The company describes it as a general-purpose multimodal generation model. Its current model card uses the more specific term “omni-modal generative system.”
Both descriptions point to the same idea. H3 can interpret several media types inside one context. It then generates a short video and native stereo audio together.
The official MiniMax H3 launch and H3 model card document these current facts:
| Item | Current specification | Scope |
|---|---|---|
| Official model name | MiniMax H3 | MiniMax launch page and model card |
| API model ID | MiniMax-H3 | Hosted Video Generation V2 API |
| Context inputs | Text, images, video, and audio | Exact combinations depend on the mode and access surface |
| Documented output | Video with native stereo audio | Hosted API and released H3-Base checkpoints |
| Duration | 4–15 seconds | Hosted/API specification |
| Frame rate | 24 fps | H3 model card |
| Audio | 32 kHz stereo | H3 model card |
| Resolution | H3-Base generates at 768p; H3 reaches up to 2K through Regenerate-2K or the end-to-end API | Do not treat local H3-Base as the complete 2K system |
| Public weights | H3-Base FL2VA and H3-Base Ref2VA | Open-weight release under a custom license |
“Omni-modal” does not mean every MiniMax product has merged into one endpoint. The documented H3 API and released checkpoints focus on joint video-and-audio output.
H3's training includes broader tasks, but that does not create separate public image-only, music-only, or audio-only H3 products. The current product surface matters more than the training-task list.
The easiest way to understand these names is to separate models from products. MiniMax is the company. H3 and M3 are different models. Hailuo AI is a creative app, while Hailuo 2.3 is a specific video model.
| Name | What it is | Primary job | Key distinction |
|---|---|---|---|
| MiniMax H3 | General-purpose omni-modal generation system | Generate video and native stereo audio from multimodal context | Official API ID: MiniMax-H3 |
| MiniMax M3 | Separate natively multimodal model | Coding, agentic work, computer use, and long-context tasks | Accepts image and video input, but its official role differs from H3 |
| Hailuo AI | User-facing MiniMax web app | Gives creators an interface for media generation | MiniMax links it as an official H3 app surface |
| Hailuo 2.3 | Specific MiniMax video model released in October 2025 | Video generation | MiniMax described it as building on Hailuo 02 |
| “Hailuo H3” or “Hailuo 3.0” | Search and community wording | Helps some users find the new release | MiniMax does not use these as H3's official model name |
This distinction prevents two common mistakes. H3 is not the video edition of M3. Hailuo 2.3 is not another official name for H3.
H3 and M3 both handle more than one input type. That shared word—multimodal—causes much of the confusion.
The practical difference lies in their jobs and outputs.
| Dimension | MiniMax H3 | MiniMax M3 |
|---|---|---|
| Primary output | Short video with native stereo audio | Responses, code, reasoning, and tool-driven task results |
| Main task scale | A 4–15-second target clip | Coding and agentic work with context windows up to one million tokens |
| Typical choice | Generated scenes, product clips, poster animation, or reference-led video | Code, analysis, computer use, or long-running agent tasks |
MiniMax's official M3 release calls M3 natively multimodal. It accepts image and video input. Describing M3 as text-only would therefore be inaccurate.
That input support does not turn M3 into the H3 video endpoint. MiniMax positions M3 around coding and agentic work. It positions H3 around audiovisual generation.
Consider a product-launch project. M3 could inspect a design brief, review UI screenshots, analyze a repository, or help coordinate tasks. H3 could use product imagery, motion references, and audio direction to generate a short clip.
The models can support the same broader project. They do not perform the same role.
Hailuo appears in both product and model naming. That makes the relationship less obvious than the H3-versus-M3 distinction.
MiniMax links Hailuo AI under the H3 repository's “Online App” section. It is a product surface where users can access MiniMax media-generation features.
That does not make Hailuo the official API name for every underlying model. Product names and model names can coexist.
MiniMax announced Hailuo 2.3 on October 28, 2025. The official Hailuo 2.3 release says it built on Hailuo 02.
MiniMax framed that update around physical movement, stylization, character micro-expressions, and motion-command response. Those descriptions come from MiniMax's own launch material.
Hailuo 2.3 and H3 have separate official names and release pages. The available sources do not say that H3 renamed or directly replaced Hailuo 2.3.
Some third-party pages and users call the new model “Hailuo H3,” “Hailuo 3.0,” or “Hailuo 03.” These terms can help identify search intent, but they should not define the product.
As of August 7, 2026, MiniMax's launch page and model card use “MiniMax H3” in prose. The Video Generation V2 API requires the model ID MiniMax-H3.
H3 supports several input patterns under one video-generation system. The right pattern depends on whether you start with an idea, a frame, or several reference assets.
With text-to-video, you describe subjects, action, setting, camera behavior, dialogue, sound effects, ambience, and music. H3 generates the visual and audio tracks together.
This differs from workflows that generate a silent clip first. It also differs from adding stock audio after generation. Native audio lets the model treat visible action and sound as parts of one target sequence.
That does not guarantee accurate dialogue, sound timing, or music on every attempt. It means the model jointly represents those output channels.
The H3-Base FL2VA checkpoint supports zero, one, or two input images.
These modes let you start from an approved composition or define a transition's endpoints. They do not guarantee that every intermediate frame will preserve all source details.
The H3-Base Ref2VA checkpoint accepts reference images, video, and audio alongside text. Each source can guide a different part of the requested result.
For example, an image can establish a character or product. A video can guide motion or camera behavior. An audio clip can guide a voice, rhythm, or sound relationship.
MiniMax's official example combines camera motion from one video, a character from an image, and vocals from an audio source. The model interprets their relationships through text.
The model card places limits on reference counts and duration. It also states that reference audio cannot serve as the only media input in local Ref2VA workflows.
H3 currently produces clips from 4 to 15 seconds. The model card documents 24 fps and 32 kHz stereo sound.
H3-Base produces 768p output. MiniMax's documented 2K workflow combines H3-Base with hosted Context-IR and Regenerate-2K APIs. The end-to-end Video Generation V2 API can also return 2K output.
The official Video Generation V2 API reference remains the best source for current request and output options.
MiniMax says its early tests showed strong instruction following, text rendering, brand rendering, and video-to-video motion transfer. These are vendor-reported claims, not independent results from this article.
The complete H3 system contains three parts. Their boundaries explain the 2K workflow and the open-weight release.
TEXT + IMAGES + VIDEO + AUDIO
│
▼
H3-Context-IR — hosted context interpretation
│
▼
H3-Base — public weights, 768p audiovisual generation
│
▼
H3-Regenerate-2K — hosted context-aware 2K regeneration
Context-IR interprets how the supplied media relates to the requested output. H3-Base then generates 768p video and stereo audio. Regenerate-2K uses that result and the original context to create the higher-resolution version.
MiniMax published H3-Base FL2VA and Ref2VA weights. It did not include Context-IR or Regenerate-2K in the current public weight release.
That makes “open-weight” more precise than “fully open source” for the complete H3 system. A custom MiniMax H3 Community License governs the downloadable materials.
H3's broad context support does not remove the normal constraints of generated video.
First, each output lasts 4 to 15 seconds. A longer ad, film, or product story still needs several shots and an editing timeline.
Second, controls depend on the access surface. An upstream model capability does not prove that every app exposes the same input roles, settings, or limits.
Third, reference inputs guide a result rather than guaranteeing exact copying. Complex motion, dialogue, small text, identities, and source relationships still need review.
Finally, MiniMax says H3 has room to improve in multimodal understanding, model scale, and visual detail. That note appears in the company's own launch roadmap.
Treat launch examples as demonstrations, not guarantees. Test your own subjects, references, text, dialogue, and motion before choosing a production workflow.
MiniMax presents H3 as a model for short-form audiovisual production. Its official examples include film titles, product pages, animated posters, advertising, and ecommerce.
The company also names branding, product design, UI/UX, and gaming as target areas. Those categories suggest several practical fits.
H3 still produces short clips. Longer stories require shot planning, repeated generation, review, and editing.
The naming is simpler than it first appears: use MiniMax H3 for the audiovisual model, MiniMax M3 for coding and agentic work, Hailuo AI for the user-facing product, and Hailuo 2.3 for the earlier video model.
MiniMax H3 is coming soon to DomoAI. We’ll share the news once it’s live—stay tuned.
Recent articles
© 2026 DOMOAI PTE. LTD.
DomoAI