
目次
MiniMax H3で良いプロンプトを書くには、まず正しいモードを選び、映像のタイムライン、会話、物理的な音、視聴者だけに聞こえる音楽を分けます。本ガイドでは、T2VA、I2VA、FL2VA、L2VA、ネイティブ音声の指示に使える独自テンプレートを紹介します。いずれもMiniMaxの公式形式に沿っており、そのままコピーして調整できます。
正しい構造はH3に意図を明確に伝える助けになりますが、カット、声、文字、動き、フレームの一致を完全に保証するものではありません。
以下のプロンプトは、MiniMaxの公式ドキュメントに掲載された例を基に作成しています。
まず、用意している映像素材を確認します。略称が高度に見えるという理由だけでモードを選ばないでください。
| モード | 入力素材 | プロンプトの最初の役割 | 適した用途 |
|---|---|---|---|
| T2VA | なし | シーン、被写体、動作、ショット、音声計画全体を設定する | テキストから新しい映像と音声のアイデアを作る |
| I2VA | 最初のフレーム1枚 | 冒頭画像を完全に参照し、その後の展開を記述する | 準備済みの冒頭構図を動かす |
| FL2VA | 最初と最後のフレーム | 両方の画像を対応させ、その間に見える変化を記述する | トランジションの始まりと終わりを制御する |
| L2VA | 最後のフレーム1枚 | 指定した最終フレームへ自然につながる前段を組み立てる | 結末から逆算した登場や変化を設計する |
| Ref2VA | 参照画像または動画と、任意の参照音声。音声だけを唯一の参照入力にはできない | 完全参照形式を使い、各素材の役割を指定する | 複数の素材から人物、スタイル、動き、カメラ、音を導く |
MiniMaxは公式H3基本プロンプトガイドで、4つの基本モードを説明しています。Ref2VAは別の6セクション構成を使うため、高度なワークフローとして分けて扱います。
基本プロンプトは、タイムライン、物理的な音、視聴者だけに聞こえる音楽という3層で考えます。
integrated_multimodal_description:
[Shot 1] Describe the subjects, setting, visible action, camera, dialogue, and any sound tied to this moment.
[Shot 2] At 00:03.000, describe the next visible and audible event.
overall_soundscape:
Describe ambience, physical sound effects, and non-verbal human sounds across the clip.
non_diegetic_music:
Describe the audience-only score, or write N/A when no score is wanted.
integrated_multimodal_descriptionには、再生順を記述します。ショット、動作、会話、歌、画面に表示する文字、同期する音をここに入れます。
overall_soundscapeには、室内音、風、交通音、足音、衝突音、衣擦れ、呼吸音などの物理的な音を記述します。会話全文をここに繰り返さないでください。
non_diegetic_musicは、視聴者だけに聞こえる音楽の欄です。シーン内のラジオや演奏者が音楽を鳴らす場合は、その出来事をタイムラインに入れます。
次の7段階で作成します。
T2VAには画像対応行がありません。3つのフィールドから始め、すべてをテキストで設定します。
以下の例では、8秒の商品登場シーンを作ります。ショット間で変えてはいけない商品記述は「a square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE"」に固定しています。
テスト前に、変化させてはいけない商品の特徴を一覧にしてください。商品動画の一貫性ガイドでは、特徴を合否基準に変える方法を説明しています。
実行条件:T2VA、8秒、16:9、入力画像なし。
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. A square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" stands on a wet black-basalt plinth inside a greenhouse before sunrise. Cold blue window light reaches the bottle from camera right. The camera makes a slow, small-amplitude push toward the plinth. Water beads slide down the glass while fern leaves move gently behind it. [Shot 2] At 00:04.000, the camera cuts to a close-up of the same square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE". A narrow amber light sweeps from left to right across the bottle. The label stays front-facing and readable. One water bead reaches the base as a quiet off-screen woman (S1) says: <d>[English] Find your north.</d> The shot holds steady through the final frame. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Light greenhouse rain taps on the glass roof. Leaves rustle softly while water drops strike the basalt surface. A low room tone continues beneath the voice.
non_diegetic_music: Three isolated felt-piano strikes sound during the first half. A bowed-glass drone rises twice, then stops two seconds before the end.
ラベル、声、光の移動、同期する水滴はタイムラインに置き、雨音と葉の音はサウンドスケープに置いています。
I2VAでは、入力画像が0.00秒時点の実際の最初のフレームになります。MiniMax指定の1行目を使い、画像ですでに定まっている内容から先の展開を書きます。
最初のフレームを文章で作り直さないでください。動きを加える前に、見えている商品、構図、照明、色、位置関係を保ちます。
実行条件:I2VA、8秒、元フレームは16:9、Picture 1が冒頭の商品ショット。
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. The square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" remains in the position, scale, lighting, and greenhouse composition established by <Picture 1>. The camera arcs clockwise with small amplitude at slow speed. Water beads move down the glass while the background fern leaves shift in a light draft. A narrow amber reflection begins at the bottle's left edge and travels across its front face. The label remains front-facing and readable. At the end, the camera stops and holds the original product as the brightest object. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Soft rain taps the greenhouse roof while leaves brush against one another. Water drops strike the basalt plinth at irregular intervals.
non_diegetic_music: A muted handpan sounds once every two seconds. One airy synthesizer tone enters halfway and recedes before the final frame.
このプロンプトは、画像から自然に変化できる内容だけを指定しています。カメラを動かしても、商品やシーンを別物に置き換えない構成です。
FL2VAは両端を固定します。プロンプトでは、Picture 1からPicture 2までに見える変化を説明する必要があります。
連続した補間を求める場合は、1つのショットを使います。カットは、制作意図として本当に必要な場合だけ追加します。
実行条件:FL2VA、8秒、同じ16:9の元フレーム、Picture 1は0.00秒、Picture 2は8.00秒。
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. At the opening instant, Picture 1 determines the square amber-glass perfume bottle's location, size, and framing. It has a matte-black cap and a cream label reading "NORTHLINE". The camera trucks right with small amplitude at slow speed. A thin stream of water crosses the basalt plinth from left to right. The greenhouse light shifts gradually from cold blue dawn toward Picture 2's warm amber direction. Fern shadows travel across the background, then settle. The bottle does not rotate, and its label remains front-facing. By the last frame, the camera, water, reflections, shadows, and lighting match Picture 2's spacing, viewing angle, and composition. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Rain decreases gradually as water runs across the stone. The final drops become farther apart before the scene settles into quiet greenhouse ambience.
non_diegetic_music: N/A
有効なFL2VAプロンプトは、カメラの移動、水の位置、光の変化、影の動きといった途中の変化を具体的に示します。
L2VAでは、Picture 1が冒頭ではなく結末になります。利用できる時間内にそのフレームへ到達できる、自然な初期状態を選びます。
公式の1行目では、最後のショット番号と小数点以下2桁の実効時間を使います。1ショット8秒のプロンプトなら、8.00-secondで終わり、Shot 1を指定します。
実行条件:L2VA、8秒、入力する最終フレームは16:9、Picture 1が最後の商品構図。
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. An empty wet black-basalt plinth fills the foreground inside a dark greenhouse before sunrise. The camera pulls out with small amplitude at slow speed. From the left, a square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" slides smoothly onto the plinth and slows as it reaches center. Cold blue light from the right changes gradually to a narrow amber beam from the left. Water beads become visible on the bottle. During the final two seconds, every moving element comes to rest. The object placement, scale, camera view, reflections, label direction, fern shadows, and illumination now match <Picture 1>. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Rain taps the greenhouse roof while glass slides softly over wet stone. The sliding sound slows and stops as one final drop strikes the plinth.
non_diegetic_music: Four muted glass harmonics sound at even intervals. The fourth ends as the bottle stops moving.
逆算する場合は、冒頭から変える要素を少数に絞ると安定しやすくなります。
最初のショットにはタイムスタンプを付けません。2つ目以降のショットは、動画の長さの範囲内で、前より後のカット時刻から始めます。
MiniMaxが示すショットラベルの形式は次のとおりです。省略記号の部分は仮の文です。
[Shot 1] Live-action, cinematic, a medium-wide shot establishes...
[Shot 2] At 00:03.500, the camera cuts to...
[Shot 3] At 00:06.250, the shot switches to...
1つのショット内で続く動作をタイムスタンプで表さないでください。カットがない場合は「ショットの中間で」など、文章で説明します。
カメラの指示は、まず動きの種類を示します。結果を明確にできる場合だけ、移動量と速度を加えます。
| 曖昧な指示 | より明確な指示 |
|---|---|
| 映画的なカメラ | カメラが低速かつ小さい移動量で前進する |
| ダイナミックな動き | カメラが高速かつ大きい移動量で左へ平行移動する |
| ラベルに注目する | 光がラベルを横切る間、カメラは静止したクローズアップを保つ |
| 商品の周囲を回る | カメラが低速かつ小さい移動量で時計回りに弧を描く |
プロンプトの構文を書く前に、カットを設計しましょう。AI動画の絵コンテガイドでは、アニメーション作成前にショットの目的と連続性を決める方法を紹介しています。
H3はネイティブステレオ音声付きの動画を生成するため、最初のプロンプトに音声の指示を含めます。4つの音の役割を区別してください。
| 音の役割 | 記述場所 | 例 |
|---|---|---|
| 会話または歌 | integrated_multimodal_description | (S1) says: <d>[English] Find your north.</d> |
| 映像に同期する音 | タイムライン内の該当ショット | 手を離すと同時にキャップが閉まる音 |
| 環境音と物理的な音 | overall_soundscape | 雨、足音、衣擦れ、衝突音、呼吸 |
| 視聴者だけに聞こえる劇伴 | non_diegetic_music | 楽器、テンポ、リズム、音量変化 |
声が初めて登場する順に(S1)、(S2)以降のIDを割り当てます。すべてのショットで、同じIDを同じ声に使います。
<d>の内側には、言語タグと実際に話す言葉だけを入れます。話者の説明、動作、話し方、IDはタグの外側に置きます。
The calm off-screen woman (S1) says: <d>[English] Find your north.</d>
The shopkeeper with a low, measured voice (S2) replies: <d>[English] It was here all along.</d>
ナレーションでは、MiniMax指定の明示的な表現を使い、画面上の人物の口を閉じたままにします。
The woman (S1) says in an off-screen voiceover: <d>[English] I kept the first bottle.</d> while her lips remain completely closed.
環境音や効果音だけを求め、視聴者向けの劇伴が不要な場合は、N/A を non_diegetic_music に書きます。全編を完全な無音にする場合だけ、overall_soundscape: N/A を使います。
画面に表示する文字は、半角のダブルクォーテーションで囲み、句読点を含めて正確に転記します。ラベルを引用符で囲むと指示は明確になりますが、文字が完全に描画される保証はありません。
スタイルを表す単語を変える前に、構造から確認します。変数は一度に1つだけ変更し、同じ合否基準で次の出力と比較してください。
| 失敗 | 考えられるプロンプトの問題 | 公式形式に沿った修正 |
|---|---|---|
| 別の人物が話す | IDが変わった、または話者が最初に定義されていない | 各声の初登場時に固定の(S1)と(S2)を割り当て、その後も同じIDを使う |
| 不要なBGMが入る | 劇伴の欄が曖昧、または欠けている | non_diegetic_music: N/A と書き、必要な物理音は overall_soundscape に残す |
| タイムスタンプが無視される | カットではなく動作に付けている、または動画の長さを超えている | 2つ目以降のショットだけに付け、動画内でカット時刻を昇順にする |
| 不規則にカットが入る | 小さな画角変更を別ショットとして書いている | 新しい情報を見せるカットでない限り、1ショット内のカメラワークとして記述する |
| FL2VAが両端の間を飛ぶ | 2つの静止画の説明を繰り返している | 最終フレームへ段階的に近づく、目に見える途中の変化を示す |
| L2VAが画像を冒頭として扱う | 最終フレームの対応行、または最後に一致させる指示がない | L2VAの指示を最初に置き、結末でのみ画像に到達させる |
| 画面上の文字が変わる | 文字を言い換えた、または引用符で囲んでいない | 正確な文字列を半角ダブルクォーテーション内に保ち、生成後に確認する |
| 会話がサウンドスケープでも繰り返される | 同じ会話を複数フィールドにコピーしている | 会話全文は映像タイムライン内の<d>だけに入れる |
これらの修正は指示を明確にしますが、すべての生成が細部まで従うことを保証するものではありません。
複数の素材から人物、動き、カメラ、声を取り入れる必要がある場合は、この基本形式を拡張し続けず、Ref2VAと公式の完全参照ガイドを使ってください。
次のチェックリストを制作メモにコピーしてください。
[Shot 1]にタイムスタンプがなく、後続のカット時刻が昇順で動画の長さ内に収まっている。(Sx)IDを保っている。<d>内にある。overall_soundscapeにある。non_diegetic_musicにあり、不要なら同欄にN/Aとある。公式のVideo Generation V2 APIリファレンスでは、空でないテキストプロンプトが必須です。また、最初・最後のフレームと参照メディアは別の役割として扱われるため、1回のAPIリクエストで両方の役割を混在させないでください。
正しいモードを選び、integrated_multimodal_description、overall_soundscape、non_diegetic_musicの3フィールドを使います。I2VA、FL2VA、L2VAでは、その前に公式の画像対応行を追加します。
T2VAはテキストから、I2VAは最初のフレームから始めます。FL2VAは最初と最後のフレームをつなぎ、L2VAは指定された最後のフレームへ到達します。モードごとに最初の指示と映像設計の向きが変わります。
声の初登場時に固定の(S1)、(S2)以降のIDを割り当て、ショット間でも同じIDを使います。言語タグと実際の台詞だけを<d>内に入れます。
同期する効果音はショットに、広い環境音はoverall_soundscapeに書きます。視聴者向けの劇伴が不要ならnon_diegetic_music: N/Aと記述します。
タイムスタンプがカットではなく動作を示している、前の時刻と同じ、または動画の長さを超えている可能性があります。Shot 2以降だけに使い、カット時刻を昇順にしてください。
H3-Base Ref2VAへ直接入力する場合は、被写体の定義、概要、保持項目、詳細な再生順、サウンドスケープ、音楽からなる6セクション形式を使います。ホスト型APIでは、参照メディアと空でないテキストプロンプトを指定できます。
これで、適切なH3モードを選び、タイムラインを組み立て、音声レイヤーを分けられます。次の生成前に対応するテンプレートを保存し、QAチェックリストを実行してください。
MiniMax H3は近日中にDomoAIへ追加される予定です。DomoAIで利用可能になり次第、お知らせします。
最近の記事
© 2025 ドメインページ(株)
どーもあい