
內容表
有效的 MiniMax H3 提示詞要先選對模式,再把視覺時間軸、對話、物理聲音與只有觀眾聽得到的音樂分開。本指南提供 T2VA、I2VA、FL2VA、L2VA 與原生音訊指示的原創範本,可直接複製後調整。每個範本都遵循 MiniMax 官方格式。
正確結構能讓 H3 更清楚地理解指示,但無法保證每次都精確符合剪接、聲音、文字、動作或畫面。
以下提示詞以 MiniMax 官方文件提供的範例為基礎。
從手上的視覺素材開始判斷。不要只因某個縮寫看起來較進階就選擇該模式。
| 模式 | 提供的輸入 | 提示詞首先要完成的事 | 適合情境 |
|---|---|---|---|
| T2VA | 無 | 建立場景、主體、動作、鏡頭與完整音訊規劃 | 從文字建立全新的影音概念 |
| I2VA | 一張首幀 | 完整參考開場圖片,再描述後續發展 | 讓已準備好的開場構圖動起來 |
| FL2VA | 首幀與尾幀 | 對齊兩張圖片,並描述兩者之間可見的變化路徑 | 控制轉場的開頭與結尾 |
| L2VA | 一張尾幀 | 設計可合理落在指定最終畫面的前段內容 | 從結尾逆推揭示或轉換 |
| Ref2VA | 參考圖片或影片,可搭配參考音訊;音訊不能是唯一參考輸入 | 透過完整參考格式分配每項素材的角色 | 用多個來源引導身分、風格、動作、鏡頭或聲音 |
MiniMax 在官方 H3 的基礎提示詞指南記載前四種基礎模式。Ref2VA 使用另一套六段式規劃格式,適合放在獨立的進階工作流程中。
可把基礎提示詞理解為三層:時間軸、物理聲音,以及只有觀眾聽得到的音樂。
integrated_multimodal_description:
[Shot 1] Describe the subjects, setting, visible action, camera, dialogue, and any sound tied to this moment.
[Shot 2] At 00:03.000, describe the next visible and audible event.
overall_soundscape:
Describe ambience, physical sound effects, and non-verbal human sounds across the clip.
non_diegetic_music:
Describe the audience-only score, or write N/A when no score is wanted.
integrated_multimodal_description承載播放順序。鏡頭、動作、對話、歌唱、畫面文字與同步聲音都寫在這裡。
overall_soundscape涵蓋室內底噪、風聲、交通聲、腳步、撞擊、衣料摩擦、呼吸等物理聲音。不要在這裡重複整段對話。
non_diegetic_music用來描述只有觀眾聽得到的配樂。若音樂來自場景中的收音機或表演者,應把該事件寫在時間軸內。
可依照以下七個步驟撰寫:
T2VA 沒有圖片對齊行。直接從三個欄位開始,所有內容都由文字建立。
以下範例生成一段 8 秒的商品揭示影片。不可在鏡頭間變動的商品描述固定為「a square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE"」。
測試前,先列出不能偏移的商品細節。商品影片一致性指南說明如何把這些細節轉成驗收標準。
執行條件:T2VA、8 秒、16:9、不提供圖片。
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. A square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" stands on a wet black-basalt plinth inside a greenhouse before sunrise. Cold blue window light reaches the bottle from camera right. The camera makes a slow, small-amplitude push toward the plinth. Water beads slide down the glass while fern leaves move gently behind it. [Shot 2] At 00:04.000, the camera cuts to a close-up of the same square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE". A narrow amber light sweeps from left to right across the bottle. The label stays front-facing and readable. One water bead reaches the base as a quiet off-screen woman (S1) says: <d>[English] Find your north.</d> The shot holds steady through the final frame. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Light greenhouse rain taps on the glass roof. Leaves rustle softly while water drops strike the basalt surface. A low room tone continues beneath the voice.
non_diegetic_music: Three isolated felt-piano strikes sound during the first half. A bowed-glass drone rises twice, then stops two seconds before the end.
標籤、聲音、光線掃動與同步水滴都放在時間軸;雨聲與葉片聲則放在overall_soundscape。
I2VA 會把提供的圖片視為 0.00 秒時的實際首幀。先使用 MiniMax 固定的第一行,再從圖片已建立的內容向後描述。
不要用文字重新設計首幀。加入動作前,先保留畫面中的商品、構圖、光線、色彩與空間關係。
執行條件:I2VA、8 秒、16:9 來源畫面、Picture 1 提供開場商品鏡頭。
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. The square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" remains in the position, scale, lighting, and greenhouse composition established by <Picture 1>. The camera arcs clockwise with small amplitude at slow speed. Water beads move down the glass while the background fern leaves shift in a light draft. A narrow amber reflection begins at the bottle's left edge and travels across its front face. The label remains front-facing and readable. At the end, the camera stops and holds the original product as the brightest object. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Soft rain taps the greenhouse roof while leaves brush against one another. Water drops strike the basalt plinth at irregular intervals.
non_diegetic_music: A muted handpan sounds once every two seconds. One airy synthesizer tone enters halfway and recedes before the final frame.
這則提示詞只描述能從圖片自然發展的變化,避免在移動鏡頭時替換商品或場景。
FL2VA 會固定兩端。提示詞必須解釋 Picture 1 到 Picture 2 之間可見的變化路徑。
若要連續插值,使用單一鏡頭即可。只有創作規劃確實需要時才加入剪接。
執行條件:FL2VA、8 秒、兩張相符的 16:9 來源畫面、Picture 1 在 0.00 秒、Picture 2 在 8.00 秒。
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. At the opening instant, Picture 1 determines the square amber-glass perfume bottle's location, size, and framing. It has a matte-black cap and a cream label reading "NORTHLINE". The camera trucks right with small amplitude at slow speed. A thin stream of water crosses the basalt plinth from left to right. The greenhouse light shifts gradually from cold blue dawn toward Picture 2's warm amber direction. Fern shadows travel across the background, then settle. The bottle does not rotate, and its label remains front-facing. By the last frame, the camera, water, reflections, shadows, and lighting match Picture 2's spacing, viewing angle, and composition. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Rain decreases gradually as water runs across the stone. The final drops become farther apart before the scene settles into quiet greenhouse ambience.
non_diegetic_music: N/A
實用的 FL2VA 提示詞會寫明中間動作,例如鏡頭位移、水的位置、光線變化與陰影移動。
L2VA 把 Picture 1 設為結尾,而不是開頭。選擇能在可用時間內合理到達該畫面的較早狀態。
官方對齊行會使用最終鏡頭編號,並以小數點後兩位表示有效片長。因此,單鏡頭 8 秒提示詞會以 8.00-second結尾,並使用Shot 1。
執行條件:L2VA、8 秒、提供的尾幀為 16:9、Picture 1 提供最終商品構圖。
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic product film. An empty wet black-basalt plinth fills the foreground inside a dark greenhouse before sunrise. The camera pulls out with small amplitude at slow speed. From the left, a square amber-glass perfume bottle, matte-black cap, cream label reading "NORTHLINE" slides smoothly onto the plinth and slows as it reaches center. Cold blue light from the right changes gradually to a narrow amber beam from the left. Water beads become visible on the bottle. During the final two seconds, every moving element comes to rest. The object placement, scale, camera view, reflections, label direction, fern shadows, and illumination now match <Picture 1>. No extra bottle, label, subtitle, or watermark appears.
overall_soundscape: Rain taps the greenhouse roof while glass slides softly over wet stone. The sliding sound slows and stops as one final drop strikes the plinth.
non_diegetic_music: Four muted glass harmonics sound at even intervals. The fourth ends as the bottle stops moving.
逆向規劃時,若開場狀態只需少量變化,通常較容易成功。
第一個鏡頭不加時間戳記。之後每個鏡頭都要在目標片長內,以嚴格遞增的剪接時間開始。
MiniMax 文件中的鏡頭標籤格式如下,省略號只是佔位內容:
[Shot 1] Live-action, cinematic, a medium-wide shot establishes...
[Shot 2] At 00:03.500, the camera cuts to...
[Shot 3] At 00:06.250, the shot switches to...
不要用時間戳記描述同一鏡頭內的連續動作。若沒有剪接,可使用「鏡頭進行到一半時」等文字表達。
運鏡指示應先寫動作類型。只有在有助於界定結果時,才加入幅度與速度。
| 模糊指示 | 更清楚的指示 |
|---|---|
| 電影感鏡頭 | 攝影機以慢速、小幅度向前推進 |
| 動態移動 | 攝影機以快速、大幅度向左平移 |
| 聚焦標籤 | 光線掃過標籤時,攝影機保持靜態特寫 |
| 環繞商品 | 攝影機以慢速、小幅度順時針繞行 |
撰寫提示詞格式前,先規劃剪接。AI 影片分鏡指南說明如何在動畫製作前定義鏡頭目的與連續性。
H3 生成帶有原生立體聲音訊的影片,因此音訊指示應直接寫入初始提示詞。請分清四種聲音任務。
| 聲音任務 | 撰寫位置 | 範例 |
|---|---|---|
| 對話或歌唱 | integrated_multimodal_description | (S1) says: <d>[English] Find your north.</d> |
| 同步事件 | 時間軸中的對應鏡頭 | 手放開時,瓶蓋同步扣上 |
| 環境聲與物理聲音 | overall_soundscape | 雨聲、腳步、衣料聲、撞擊與呼吸 |
| 只有觀眾聽到的配樂 | non_diegetic_music | 樂器、速度、節奏與音量變化 |
依聲音首次出現的順序分配(S1)、(S2)與後續 ID。所有鏡頭中,同一個 ID 都必須對應同一個聲音。
<d>內只放語言標籤與實際台詞。說話者描述、動作、語氣與 ID 要留在標籤外。
The calm off-screen woman (S1) says: <d>[English] Find your north.</d>
The shopkeeper with a low, measured voice (S2) replies: <d>[English] It was here all along.</d>
使用旁白時,採用 MiniMax 的明確說法,並讓畫面人物保持嘴巴閉合:
The woman (S1) says in an off-screen voiceover: <d>[English] I kept the first bottle.</d> while her lips remain completely closed.
若只要環境聲與音效、不需要觀眾聽到的配樂,請在non_diegetic_music下寫N/A。只有要求全片完全靜音時,才使用overall_soundscape: N/A。
畫面文字需要使用半形雙引號,並原樣保留標點。用引號標示標籤能讓指示更清楚,但不保證文字一定完全正確。
先檢查結構,再更換風格字詞。每次只改一個變數,並用同一套驗收標準比較下一次輸出。
| 問題 | 可能的提示詞原因 | 符合官方格式的修正 |
|---|---|---|
| 錯誤人物說話 | ID 改變,或未先建立說話者 | 每個聲音初次出現時分配固定的(S1)與(S2),之後持續沿用 |
| 出現不需要的背景音樂 | 配樂欄位模糊或遺漏 | 把需要的物理聲音保留在overall_soundscape,並寫上non_diegetic_music: N/A |
| 時間戳記似乎被忽略 | 時間戳記標示動作而非剪接,或超出片長 | 只標示第二個及後續鏡頭,並讓剪接時間在片長內嚴格遞增 |
| 影片隨機剪接 | 把微小角度變化寫成不同鏡頭 | 除非剪接要帶入新資訊,否則保留單一鏡頭並描述運鏡 |
| FL2VA 在兩端之間跳接 | 提示詞只重複兩個靜態描述 | 寫出逐步到達最終畫面的可見中間變化 |
| L2VA 把圖片當成開場 | 缺少尾幀對齊行或最後收斂的指示 | 把 L2VA 指示放在最前面,只在結尾到達指定圖片 |
| 畫面文字偏移 | 文字遭改寫或未加引號 | 把確切文字放在半形雙引號內,生成後再檢查 |
| 對話在 soundscape 重複 | 台詞被複製到多個欄位 | 完整對話只放在視覺時間軸的<d>內 |
這些修正能提升指示清晰度,但不代表每次生成都會遵循所有細節。
若提示詞要從多個來源檔案取用身分、動作、鏡頭與聲音,不要繼續擴充基礎格式。改用 Ref2VA 與官方完整參考指南。
把以下清單複製到製作筆記中:
[Shot 1]沒有時間戳記;後續剪接時間遞增且未超出片長。(Sx)ID。<d>內。overall_soundscape。non_diegetic_music,不需要時該欄寫N/A。官方Video Generation V2 API 參考文件要求非空白文字提示詞。它也把首尾幀角色與參考媒體角色分開,因此不要在同一個 API 請求中混用兩類角色。
先選對模式,再使用 integrated_multimodal_description、overall_soundscape 與 non_diegetic_music 三個欄位。I2VA、FL2VA 或 L2VA 還要先加入官方圖片對齊行。
T2VA 從文字開始,I2VA 從首幀開始,FL2VA 連接首尾幀,L2VA 則落在提供的尾幀。每種模式的第一項指示與視覺規劃方向都不同。
聲音初次出現時分配固定的 (S1)、(S2) 與後續 ID,並在所有鏡頭中沿用。<d> 內只放語言標籤與實際台詞。
同步音效寫在對應鏡頭,較廣的環境聲寫在 overall_soundscape。接著寫 non_diegetic_music: N/A,表示不需要只有觀眾聽得到的配樂。
時間戳記可能標示了動作而非剪接、與較早時間重複,或超出片長。只在 Shot 2 以後使用,並讓剪接時間嚴格遞增。
直接為 H3-Base Ref2VA 撰寫時,請使用主體定義、摘要、保留項目、詳細播放過程、overall_soundscape 與音樂的六段格式。託管式 API 可接收非空白提示詞與參考媒體。
現在可以依輸入選擇正確的 H3 模式、撰寫時間軸,並分開各個音訊層。儲存對應範本,下次生成前先完成 QA 檢查。
MiniMax H3 即將加入 DomoAI;可在 DomoAI 使用後,我們會立即公布消息。
最近的文章