
Table of Content

Try DomoAI, the Best AI Animation Generator
Turn any text, image, or video into anime, realistic, or artistic videos. Over 30 unique styles available.
MiniMax H3 reference-to-video works best when every source has a named job. Use a picture for appearance or composition, video for motion or camera, and audio for a sound signal. Then state what H3 should retain, transfer, change, or treat as a loose reference.
One asset can provide several subjects or attributes. The “one job” rule sets one primary responsibility for each source. It is a planning method, not a technical one-asset limit.
Reference-to-video, or Ref2VA, uses text plus reference images, videos, or audio to guide a new audiovisual clip. The sources can describe identity, appearance, style, movement, camera, timing, voice, sound effects, or music.
H3 does not treat every file as a frame. A picture can define appearance, video can guide camera, and audio can guide timbre.
MiniMax's official full-reference prompt guide formalizes those relationships. It uses labels, retention markers, and a six-section planning format.
That format describes intended relationships. It does not guarantee exact identity, physics, lip sync, timing, text, or pixel-level reproduction.
Choose FL2VA when a supplied image must occupy the first frame, last frame, or both. Choose Ref2VA when sources should guide attributes or behavior without serving as fixed endpoints.
| Need | Recommended H3 mode | Why |
|---|---|---|
| Start from one opening image | FL2VA checkpoint in first-frame I2VA mode | The supplied image anchors the opening state |
| Land on one supplied ending image | FL2VA checkpoint in L2VA mode | The prompt builds a path to the final state |
| Constrain both first and last images | FL2VA checkpoint in first-and-last-frame mode | Both endpoints are explicit |
| Use identity, style, motion, camera, or sound from several sources | Ref2VA | Each source can receive a different reference role |
| Mix first/last-frame API roles with reference roles in one request | Not supported in a single API request | The official API treats these role families as mutually exclusive |
The Video Generation V2 API reference separates endpoint-frame roles from reference-media roles. One API request cannot mix those role families.
Every reference adds another relationship that H3 must interpret.
Use this five-check filter before upload:
A clean packshot usually guides geometry better than a busy collage. A stable movement clip communicates camera intent better than a fast montage.
Plan the shot before gathering references. This AI storyboard guide helps separate composition, action, and camera decisions before animation.
MiniMax's current H3 repository and model documentation list these local Ref2VA limits:
The hosted API requires a non-empty text prompt in every request. Its current request body limit is 64 MB.
API reference images can use JPG, JPEG, PNG, WEBP, HEIC, or HEIF. Reference video accepts MP4 or MOV, while reference audio accepts WAV or MP3.
API upload requirements can change, so check the current Video Generation V2 API reference before submitting a request.
MiniMax lists ComfyUI as one local inference option. Local input rules and hosted API request rules may differ.
Create a Reference Job Sheet before writing the prompt. Give each file one primary job and, when needed, one secondary job.
| Source or label | Good primary jobs | What it does not guarantee |
|---|---|---|
<Subject N> | Reusable person, object, environment, action, pose, effect, or style abstracted from sources | Perfect identity lock or a separate uploaded file |
<Picture N> | Concrete frame, composition, storyboard anchor, spatial arrangement, or lighting anchor | Complete motion, temporal structure, or perfect geometry retention |
<Video N> | Source edit, continuation, motion pattern, camera path, cuts, rhythm, or temporal structure | Guaranteed replacement, physics, or pixel-level copying |
<Audio N> | Voice timbre, beat, dialogue signal, sound effect, music traits, continuity, or direct audio reuse | Guaranteed lip sync or audio-only local Ref2VA input |
Here is the worksheet format:
| Asset ID | Type | Rights cleared? | Primary job | Secondary job | Retain, transfer, or change | Shot range | Conflict or acceptance check |
|---|---|---|---|---|---|---|---|
| Picture 1 | Image | Yes | Bottle appearance and label | Material color | Retain geometry and matte-black cap | All shots | Label remains front-facing |
| Picture 2 | Image | Yes | Greenhouse environment | Dawn palette | Transfer setting, not its props | All shots | No second product appears |
| Video 1 | Video | Yes | Clockwise camera arc | Slow pacing | Transfer camera path only | Shot 1 | Bottle does not inherit source shape |
| Audio 1 | Audio | Yes | Glass-chime texture | Final accent timing | Reference texture, do not copy signal | Shot 2 | One clean chime after cap click |
Acceptance checks make each role reviewable.
For product work, the product-video consistency guide adds geometry and label checks.
This direction has no priority:
Use the silver cap from Picture 1 and the black cap from Picture 2. Keep both exact.
This version assigns responsibility:
Picture 1 controls the bottle geometry, amber glass, matte-black cap, and label. Picture 2 controls only the greenhouse environment and dawn lighting; ignore the container shown in Picture 2.
When a reference performance matters more than its performer, state that boundary explicitly: Video 1 controls motion and camera only; ignore its performer and setting.
Keep label numbers stable across all six sections.
<Subject N> means reusable or editable visible content, such as a person, product, environment, action, pose, interface, or style.
<Subject 1> is the square amber-glass perfume bottle whose geometry, matte-black cap, and cream "NORTHLINE" label come from <Picture 1>.
Use standalone <Picture N> when the image acts as a concrete frame or composition anchor.
<Picture 3> is the first-frame composition anchor for [Shot 1], defining the bottle position, plinth angle, and empty space on the right.
If an image only defines a subject, cite it inside the Subject definition.
Use <Video N> for source editing, continuation, camera, cuts, rhythm, or temporal structure. Extracted people and objects still need Subject labels.
Use <Audio N> for direct reuse or reference to voice, timing, beat, effects, or music traits. Audio and Video numbers remain independent.
MiniMax's full-reference guide uses six sections in this order:
subject_definitionssummaryretention_analysisdetailed_descriptionoverall_soundscapenon_diegetic_musicThe hosted API requires text, but it does not require every user to hand-author this full rewrite. H3-Context-IR can process free-form multimodal input. Use the six sections as a precise planning language, not a mandatory raw API schema.
The original example below applies MiniMax's documented format to a fictional product shot.
Example run card: Ref2VA · 8 seconds · 16:9 · three images · one video · one audio reference.
subject_definitions:
<Subject 1> is the square amber-glass perfume bottle whose geometry, matte-black cap, and cream "NORTHLINE" label come from <Picture 1>.
<Subject 2> is the greenhouse environment from <Picture 2>, with wet black-basalt surfaces, tall fern leaves, glass roof panels, and cold blue dawn light.
<Picture 3> is the first-frame composition anchor for [Shot 1], defining the bottle position, plinth angle, and empty space on the right.
<Video 1> provides the slow clockwise camera arc and measured pacing for [Shot 1], without providing product appearance.
<Audio 1> is the sound-texture reference for one clear glass chime heard after the cap click in [Shot 2].
summary:
[keyframe completion + reference generation + audio reference] The target video presents <Subject 1> inside <Subject 2>. It begins from the composition in <Picture 3>, follows the camera path from <Video 1>, and uses <Audio 1> only as a sound-texture reference for the final chime.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the square amber-glass body, matte-black cap, cream "NORTHLINE" label, and front-facing product identity are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the wet basalt, fern leaves, glass roof, and cold blue dawn lighting are retained.
<Picture 3> ([Shot 1] first frame): fully_preserved - the opening bottle position, plinth angle, and right-side negative space are retained.
<Video 1> (camera path and pacing): partially_preserved - its slow clockwise arc and measured speed guide [Shot 1], while its source subjects and setting are excluded.
<Audio 1>: reference - its clean glass-chime texture guides one newly generated accent without copying the original signal.
detailed_description:
The target video uses a live-action cinematic product-film style with restrained movement and cool dawn lighting.
[Shot 1] The shot begins from <Picture 3>. <Subject 1>, the square amber-glass perfume bottle with a matte-black cap and cream "NORTHLINE" label, stands on the wet basalt plinth inside <Subject 2>. The camera follows the slow clockwise arc and measured pacing referenced from <Video 1>. Cold blue light enters from the right through the greenhouse roof. Water beads slide down the amber glass. Fern leaves move slightly behind the product. The bottle stays fixed, front-facing, and unchanged while the camera completes the arc. No second bottle, subtitle, or watermark appears.
[Shot 2] At 00:05.000, the camera cuts to a static close-up of <Subject 1>. A clean hand enters from the upper right, presses the matte-black cap down once, and releases it. The cap makes a dry click. A single newly generated glass chime follows, using the clear texture referenced from <Audio 1> without copying its signal. A narrow amber reflection crosses the cream "NORTHLINE" label from left to right. The hand exits. The bottle remains centered and readable through the final frame.
overall_soundscape:
Light rain taps on the glass roof throughout the video. Fern leaves brush softly in the background, while water drops strike the basalt and the cap produces one dry mechanical click.
non_diegetic_music:
N/A
The plan states what each source controls and excludes.
Assign roles by shot because a reference may matter for only part of the clip.
Use this order for each shot:
The first shot has no timestamp. Later shots use [Shot N] At MM:SS.mmm, ... with increasing cut times. If two videos guide movement, assign body motion to one and camera to the other.
Choose the audio marker that matches the intended signal relationship.
| Marker | Plain-language meaning |
|---|---|
fully_copy | Reuse the entire source audio as the complete final audio track |
partially_copy | Copy part of the timeline or selected layers, with changes elsewhere |
reference | Generate new audio guided by timbre, rhythm, style, words, beat, or sound texture |
weak_reference | Keep only broad category or atmospheric similarity |
When audio guides a target speaker, bind it to the same stable (Sx) speaker ID used in detailed_description. If only the timbre transfers, do not copy the source dialogue into the target.
Remember the local restriction: audio cannot be the only media reference for H3-Base-Ref2VA. Include an image or video.
Score each intended job as kept, partly kept, missing, or conflicted.
| Check | Question |
|---|---|
| Identity | Does the intended person or object remain recognizable across its shots? |
| Appearance | Are the named colors, clothing, material, label, or style present? |
| Geometry | Did the product keep its defining shape and part relationships? |
| Motion | Does the requested action follow the intended path and pace? |
| Camera | Did the camera use the requested direction without inheriting unwanted source content? |
| Timing | Did each reference begin and stop in the intended shot range? |
| Sound | Was audio copied, referenced, or newly generated as planned? |
| Conflicts | Did two references compete for the same attribute? |
Revise one relationship at a time. If only the camera failed, clarify Video 1's camera-only role and exclude its source product.
Check the official inputs and your role plan before blaming model behavior.
Cache, quantization, step count, nodes, and VRAM can affect a specific local setup. Their effects vary by configuration.
Local H3-Base produces 768p video with native stereo audio. The public release includes separate FL2VA and Ref2VA checkpoints.
The complete system also includes hosted H3-Context-IR and H3-Regenerate-2K. Regenerate-2K uses the 768p result plus the original context to regenerate at up to 2K.
Regenerate-2K is not a conventional upscaler, and it is not in the current open-weight release. Downloading Ref2VA alone therefore does not provide the complete local 2K workflow.
FL2VA anchors a first frame, last frame, or both. Ref2VA assigns identity, style, motion, camera, timing, or sound roles to reference images, video, and audio.
The local model card lists up to 9 images, 3 videos, and 3 audio clips, with no more than 12 mixed files. Video and audio duration totals each cannot exceed 15 seconds.
Not in the documented local Ref2VA workflow. Reference audio must accompany an image or video. The hosted API follows a separate set of request rules.
No. The API requires text, but users do not have to submit all six sections manually. The structure is a precise planning method.
Video can guide a motion pattern, camera path, rhythm, or temporal structure. It should be treated as guidance, not guaranteed reproduction.
H3-Base generates 768p locally. The official 2K path depends on the hosted Regenerate-2K stage, which is not included in the current open-weight release.
Build the Reference Job Sheet before your next Ref2VA request. Give each source one primary job, then review the first output against that job.
MiniMax H3 is coming soon to DomoAI. We’ll share the news when it becomes available there.
Recent articles
© 2026 DOMOAI PTE. LTD.
DomoAI