
Table of Content

Try DomoAI, the Best AI Animation Generator
Turn any text, image, or video into anime, realistic, or artistic videos. Over 30 unique styles available.
Good MiniMax H3 prompts start with the correct mode. They then separate the visual timeline, dialogue, physical sound, and audience-only music. This guide gives original, copyable templates for T2VA, I2VA, FL2VA, L2VA, and native-audio direction. Each template follows MiniMax's official format.
A correct structure gives H3 clearer instructions, but it cannot guarantee exact cuts, voices, text, motion, or frame matching.
The prompts below are based on examples provided in MiniMax's official documentation.
Start with your supplied visual material. Do not choose a mode because its acronym sounds more advanced.
| Mode | Supplied input | Prompt's first job | Best fit |
|---|---|---|---|
| T2VA | None | Establish the scene, subjects, action, shots, and complete audio plan | Create a new audiovisual idea from text |
| I2VA | One first frame | Fully reference the opening image, then describe what happens next | Animate a prepared opening composition |
| FL2VA | First and last frames | Align both pictures and describe an observable path between them | Control the opening and ending of a transition |
| L2VA | One last frame | Build a plausible earlier sequence that lands on the supplied final frame | Design an ending-first reveal or transformation |
| Ref2VA | Reference images and/or video, with optional reference audio; audio cannot be the sole reference input | Assign each asset a role through the full-reference format | Guide identity, style, motion, camera, or sound with several sources |
MiniMax documents these four base modes in its official H3 base prompt guide. Ref2VA uses a different six-section planning format, so it belongs in a separate advanced workflow.
Think of the base prompt as three layers: timeline, physical sound, and audience-only music.
integrated_multimodal_description:
[Shot 1] Describe the subjects, setting, visible action, camera, dialogue, and any sound tied to this moment.
[Shot 2] At 00:03.000, describe the next visible and audible event.
overall_soundscape:
Describe ambience, physical sound effects, and non-verbal human sounds across the clip.
non_diegetic_music:
Describe the audience-only score, or write N/A when no score is wanted.
integrated_multimodal_description carries the playback order. Put shots, action, dialogue, singing, visible text, and synchronized sound inside it.
overall_soundscape covers room tone, wind, traffic, footsteps, impacts, fabric, breathing, or similar physical sound. Do not repeat complete dialogue here.
non_diegetic_music covers music that only the audience hears. If a radio or performer creates music inside the scene, keep that event in the timeline instead.
Use this seven-part workflow:
T2VA has no picture-alignment line. Begin with the three fields and establish everything from text.
The example below creates an eight-second product reveal. Its locked product clause stays identical across shots: “a square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE".”
Before testing, list the product details that cannot drift. The product-video consistency guide shows how to turn those details into an acceptance check.
Run card: T2VA · 8 seconds · 16:9 · no supplied image.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. A square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" stands on a wet black-basalt plinth inside a greenhouse before sunrise. Cold blue window light reaches the bottle from camera right. The camera makes a slow, small-amplitude push toward the plinth. Water beads slide down the glass while fern leaves move gently behind it. [Shot 2] At 00:04.000, the camera cuts to a close-up of the same square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE". A narrow amber light sweeps from left to right across the bottle. The label stays front-facing and readable. One water bead reaches the base as a quiet off-screen woman (S1) says: <d>[English] Find your north.</d> The shot holds steady through the final frame. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Light greenhouse rain taps on the glass roof. Leaves rustle softly while water drops strike the basalt surface. A low room tone continues beneath the voice.
non_diegetic_music: Three isolated felt-piano strikes sound during the first half. A bowed-glass drone rises twice, then stops two seconds before the end.
The label, voice, light sweep, and synchronized water drop sit in the timeline. Rain and leaf movement sit in the soundscape.
I2VA treats the supplied image as the actual first frame at 0.00 seconds. Use MiniMax's fixed first line, then write forward from what the picture already establishes.
Do not redesign the first frame in prose. Preserve its visible product, composition, lighting, colors, and spatial relationships before adding motion.
Run card: I2VA · 8 seconds · 16:9 source frame · Picture 1 supplies the opening product shot.
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. The square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" remains in the position, scale, lighting, and greenhouse composition established by <Picture 1>. The camera arcs clockwise with small amplitude at slow speed. Water beads move down the glass while the background fern leaves shift in a light draft. A narrow amber reflection begins at the bottle's left edge and travels across its front face. The label remains front-facing and readable. At the end, the camera stops and holds the original product as the brightest object. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Soft rain taps the greenhouse roof while leaves brush against one another. Water drops strike the basalt plinth at irregular intervals.
non_diegetic_music: A muted handpan sounds once every two seconds. One airy synthesizer tone enters halfway and recedes before the final frame.
The prompt names only changes that can develop from the image. It avoids replacing the product and scene while moving the camera.
FL2VA anchors both ends. The prompt must explain the visible path between Picture 1 and Picture 2.
Use one shot when you want continuous interpolation. Add cuts only when the creative plan truly requires them.
Run card: FL2VA · 8 seconds · matching 16:9 source frames · Picture 1 at 0.00 · Picture 2 at 8.00.
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. At the opening instant, Picture 1 determines the square amber-glass perfume bottle's location, size, and framing. It has a matte-black cap and a cream label reading "NORTHLINE". The camera trucks right with small amplitude at slow speed. A thin stream of water crosses the basalt plinth from left to right. The greenhouse light shifts gradually from cold blue dawn toward Picture 2's warm amber direction. Fern shadows travel across the background, then settle. The bottle does not rotate, and its label remains front-facing. By the last frame, the camera, water, reflections, shadows, and lighting match Picture 2's spacing, viewing angle, and composition. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Rain decreases gradually as water runs across the stone. The final drops become farther apart before the scene settles into quiet greenhouse ambience.
non_diegetic_music: N/A
A useful FL2VA prompt names the intermediate motion: camera travel, water position, light change, and shadow movement.
L2VA makes Picture 1 the ending, not the opening. Choose a plausible earlier state that can reach that frame within the available time.
The official line uses the final shot number and an effective duration with exactly two decimal places. A single-shot eight-second prompt therefore ends at 8.00-second and uses Shot 1.
Run card: L2VA · 8 seconds · 16:9 supplied final frame · Picture 1 supplies the final product composition.
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. An empty wet black-basalt plinth fills the foreground inside a dark greenhouse before sunrise. The camera pulls out with small amplitude at slow speed. From the left, a square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" slides smoothly onto the plinth and slows as it reaches center. Cold blue light from the right changes gradually to a narrow amber beam from the left. Water beads become visible on the bottle. During the final two seconds, every moving element comes to rest. The object placement, scale, camera view, reflections, label direction, fern shadows, and illumination now match <Picture 1>. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Rain taps the greenhouse roof while glass slides softly over wet stone. The sliding sound slows and stops as one final drop strikes the plinth.
non_diegetic_music: Four muted glass harmonics sound at even intervals. The fourth ends as the bottle stops moving.
Backward planning works best when the opening state requires only a few changes.
The first shot has no timestamp. Every later shot starts with a strictly increasing cut time inside the target duration.
MiniMax documents this shot-label pattern; the ellipses below are placeholders:
[Shot 1] Live-action, cinematic, a medium-wide shot establishes...
[Shot 2] At 00:03.500, the camera cuts to...
[Shot 3] At 00:06.250, the shot switches to...
Do not use a timestamp to describe continuous action inside one shot. Use prose such as “halfway through the shot” when a cut does not occur.
Camera direction needs a motion type first. Add amplitude and speed only when they help define the result.
| Weak direction | Clearer direction |
|---|---|
| Cinematic camera | The camera pushes in with small amplitude at slow speed |
| Dynamic movement | The camera trucks left with large amplitude at fast speed |
| Focus on the label | The camera holds a static close-up while the light crosses the label |
| Circle the product | The camera arcs clockwise with small amplitude at slow speed |
Plan cuts before you write prompt syntax. The current AI storyboard guide explains how to define shot purpose and continuity before animation.
H3 generates video with native stereo audio, so audio direction belongs in the initial prompt. Keep four sound jobs distinct.
| Sound job | Where to write it | Example |
|---|---|---|
| Dialogue or singing | integrated_multimodal_description | (S1) says: |
| Synchronized event | Current shot in the timeline | The cap clicks shut as the hand releases it |
| Ambience and physical sound | overall_soundscape | Rain, footsteps, cloth, impacts, breathing |
| Audience-only score | non_diegetic_music | Instrumentation, tempo, rhythm, and volume change |
Assign (S1), (S2), and later IDs in the order voices first occur. Keep each ID attached to the same voice across all shots.
Put only the language tag and exact spoken words inside <d>. Keep the speaker description, action, delivery, and ID outside the tag.
The calm off-screen woman (S1) says: <d>[English] Find your north.</d>
The shopkeeper with a low, measured voice (S2) replies: <d>[English] It was here all along.</d>
For voiceover, use MiniMax's explicit phrase and keep the on-screen mouth closed:
The woman (S1) says in an off-screen voiceover: <d>[English] I kept the first bottle.</d> while her lips remain completely closed.
Write N/A under non_diegetic_music when you want ambience and effects without an audience-only score. Use overall_soundscape: N/A only when you request complete silence throughout.
Visible text needs double quotation marks. Copy it exactly, including punctuation. Quoting a label can improve instruction clarity, but it does not guarantee perfect text rendering.
Start with structure before changing style words. Change one variable at a time, then compare the next output with the same acceptance check.
| Failure | Likely prompt problem | Official-format correction |
|---|---|---|
| The wrong person speaks | IDs changed or a speaker was never established | Assign stable (S1) and (S2) IDs at each voice's first event, then reuse them |
| Unwanted background music appears | The score field was vague or missing | Write non_diegetic_music: N/A while keeping wanted physical sounds in overall_soundscape |
| A timestamp seems ignored | The timestamp marks an action, not a cut, or falls outside the duration | Timestamp only later shots and keep cut times strictly increasing within the clip |
| The video cuts randomly | Minor angle changes were written as separate shots | Keep one shot and describe camera motion unless a cut introduces new information |
| FL2VA jumps between endpoints | The prompt repeats two static descriptions | Name observable intermediate changes that progressively reach the final frame |
| L2VA treats the image as an opener | The last-frame alignment line or final convergence is missing | Put the L2VA instruction first and land on the picture only at the ending |
| Visible text drifts | Text was paraphrased or left unquoted | Preserve the exact text inside English double quotation marks, then review the result |
| Dialogue repeats in the soundscape | Spoken words were copied into multiple fields | Keep complete dialogue only inside in the visual timeline |
These corrections improve instruction clarity. They do not prove that every generation will follow every detail.
If a prompt needs identity, motion, camera, and voice from several source files, stop expanding this base format. Use Ref2VA and the official full-reference guide.
Copy this checklist into your production notes:
[Shot 1] has no timestamp; later cut times increase and stay inside the duration.(Sx) ID.<d>.overall_soundscape.non_diegetic_music, or that field says N/A.The official Video Generation V2 API reference requires a non-empty text prompt. It also separates first/last-frame roles from reference-media roles, so do not mix those role families in one API request.
The best base structure starts with the correct mode and uses three fields: integrated_multimodal_description, overall_soundscape, and non_diegetic_music. Add the official picture-alignment line before those fields for I2VA, FL2VA, or L2VA.
T2VA starts from text, I2VA starts from a first frame, FL2VA connects first and last frames, and L2VA lands on a supplied last frame. Each mode changes the first prompt instruction and the direction of the visual plan.
Assign stable (S1), (S2), and later IDs when voices first occur. Reuse each ID across shots, and place only the language tag plus spoken words inside <d>.
Describe synchronized effects in the shot and broader ambience in overall_soundscape. Then write non_diegetic_music: N/A so you are not requesting an audience-only score.
A timestamp may fail when it marks an action rather than a cut, repeats an earlier time, or falls outside the clip. Use timestamps only for Shot 2 and later, with strictly increasing cut times.
For direct H3-Base Ref2VA prompting, use the six-section format for subject definitions, summary, retention, detailed playback, soundscape, and music. The hosted API accepts a non-empty text prompt with reference media, so API callers do not necessarily hand-author the full rewrite.
You can now choose the right H3 mode, write its timeline, and separate each audio layer. Save the matching template and run the QA checklist before your next generation.
MiniMax H3 is coming soon to DomoAI. We’ll share the news once it’s live—stay tuned.
Recent articles
© 2026 DOMOAI PTE. LTD.
DomoAI