
Table of Content

Try DomoAI, the Best AI Animation Generator
Turn any text, image, or video into anime, realistic, or artistic videos. Over 30 unique styles available.
A short film is not one generation. It is a shot list of separate 4-to-30-second clips, generated in DomoAI Omni Reference and cut together in an editor. Your unit of work is the shot, never the film. Once you plan that way, everything else on this page follows.
Seedance 2.5 is ByteDance's model, integrated by DomoAI in Omni Reference. It generates from 4 to 30 seconds at 480P or 720P. There is no setting that turns those clips into a finished piece for you.
Do the arithmetic before you do anything else.
At 4 to 30 seconds per generation, a 90-second film is roughly 6 to 15 shots. Most pieces land nearer the top of that range, because short shots carry cuts better than long ones. Then add re-rolls.
Re-rolls are not a sign you did it wrong. They are the normal cost of the medium. Budget a few per shot and the schedule stops feeling like failure.
Two more things belong in the plan before you start. Pick one aspect ratio and hold it across every shot. Pick one resolution and do the same. Changing either halfway means regenerating everything that came before.
The person who put this best spent six to eight months on a single retro trailer. In their words: "AI makes it fairly easy to generate an interesting isolated shot. Getting dozens of shots — made with different models, months apart — to feel like they came from the same movie was the real challenge."
That gap is what this page is about.
Work in four columns: beat, shot, duration, reference set.
Give every shot one action and one camera move. Not two. Crowded prompts collapse toward the average, and creators report it consistently — one tester's verdict was that "the more you add to it, the more it changes towards some generic anime stuff."
Write durations as numbers, not adjectives. "Six seconds" is a plan. "A short beat" is not.
The list itself is four columns and nothing more:
| Beat | Shot | Duration | References |
|---|---|---|---|
| she finds the tool | lead at the workbench, back to camera | 8s | workshop, identity |
| she recognises it | tight on hands | 6s | workshop, identity, tool |
| she decides | over-shoulder to the doorway | 12s | workshop, identity |
Fill it in before you generate anything. A shot list you can read in one screen is a shot list you can budget.
If you want help turning a script into shots first, we wrote the planning half separately: storyboard your AI video.
Every shot draws from the same small set of assets. Build them in this order.
Never generate the first identity image in GPT Image 2. Every variant inherits from the anchor, so the anchor has to be the strongest version of that face you will get.
In Omni Reference, references are addressed by upload order. Image 1 is the first image you uploaded. Give each one exactly one job: identity, opening frame, prop, motion reference, audio.
Use the minimum sufficient set. The surface accepts 30 images, 10 videos and 10 audio files. Extra references without a named role only add ambiguity.
Midjourney — identity anchor
an original adult film lead in their thirties with a distinctive scarred left
eyebrow and steady, guarded expression, full-body front-facing identity portrait,
worn charcoal work jacket and dark trousers, plain mid-grey studio backdrop,
even softbox lighting, realistic skin and fabric texture, clean character
reference, generous negative space --ar 16:9 --style raw --s 70
--no text, logo, celebrity, copyrighted character, extra person, watermarkThis is the section that changes how you plan.
A cut re-anchors the model on your reference images. Identity, wardrobe and props reset from Image 1 at the start of every generation. That reset is a feature, and it is why a cut protects you.
What does not survive a cut: exact pose geometry, exact prop position, continuous motion. A hand that ended raised will not start raised in the next clip.
So do not write end-state handoffs between shots. Do not tell shot 4 to continue shot 3's body position. A new angle reads as a new setup, and forcing centimetre-level continuity across a cut is the fastest way to break everything downstream.
Plan continuity at the level of who, where and what, not at the level of limbs.
Set your expectations here, because no competing page will set them for you.
In a 152-take same-prompt comparison run in ComfyUI, one tester logged this: "Prompted a wide shot, got a medium; prompted a tight insert on hands, got a wide. Only 9 of 36 shots matched the H3 framing." The same run recorded a refusal on head count — a shot needing exactly three clothed men produced four, two shirtless, after eight takes and two prompt rewrites.
Those are other models on another surface, so read them as what creators report rather than as Seedance 2.5 behaviour. The planning lesson holds regardless.
Two things follow. Budget re-rolls per shot in your schedule. And lock counts positively — write "Exactly two people appear in this video", not "no extra people".
One lead, two locations, a turn at shot 5.
| # | Duration | References | The one thing this shot must hold |
|---|---|---|---|
| 1 | 8s | Image 1 (workshop), Image 2 (identity) | the room's light direction |
| 2 | 6s | same | the lead's jacket |
| 3 | 12s | same + Image 3 (the tool) | the tool's shape |
| 4 | 10s | same | the lead's position relative to the bench |
| 5 | 8s | Image 1 (street), Image 2 (identity) | identity only — new location |
| 6 | 12s | same | the street's light direction |
| 7 | 10s | same | the jacket again |
| 8 | 14s | same + Image 3 | the tool, now outdoors |
| 9 | 10s | same | the lead's eyeline to camera |
Shot 5 is the teaching moment. The location reference changes completely, and the identity reference does not. That is re-anchoring, and it is what lets a film move between spaces without the lead becoming somebody else.
Here is shot 1 in full.
Seedance 2.5 — Omni Reference — shot 1
Image 1 controls the workshop location, camera position, light direction and
composition, and is the literal first frame.
Image 2 controls the lead's facial identity, hairstyle and wardrobe ONLY. Do not
recreate its framing, composition, or lighting.
Create an 8-second 16:9 shot. The lead stands at the workbench with their back
three-quarters to camera, sets down a metal tool, and turns their head toward the
doorway. The camera holds a slow push in from bench height. Warm tungsten key from
the left with a hard falloff into the room's depth. Film stock, restrained grain,
visible dust in the light shaft.
Preserve the exact face, age, hairstyle, skin tone and body proportions. Exactly
one person appears in this video. No readable text, logos, signage or subtitles.And here is shot 5, where the reference set changes.
Seedance 2.5 — Omni Reference — shot 5
Image 1 controls the night street location, camera position, light direction and
composition, and is the literal first frame.
Image 2 controls the lead's facial identity, hairstyle and wardrobe ONLY. Do not
recreate its framing, composition, or lighting.
Create an 8-second 16:9 shot. The lead steps off the kerb into the road and stops,
looking back the way they came. The camera tracks with them at shoulder height and
settles as they stop. Cool sodium streetlight from above and behind, hard shadows
on wet asphalt. Film stock, restrained grain.
Preserve the exact face, age, hairstyle, skin tone and body proportions from
Image 2. Exactly one person appears in this video. No readable text, logos,
signage or subtitles.Notice what stayed identical: the identity clause, the head-count lock and the text ban. Notice what changed: everything about the room. That is the whole method.
An audio reference in Omni Reference controls soundtrack or dialogue timing. It does not detect beats. It does not place cuts.
Cut timing happens in your editor, always. If your piece is music-led and the cuts need to land on the beat, that arithmetic lives on the 2000s look recipe. It works the same way for any genre.
Everything that should feel consistent across the whole film happens after the cut, not inside the generations.
Grade once, over the finished timeline. Add grain, texture and any period treatment as a single global layer. A generated artifact rides on the subject and deforms with it; a post layer stays put on the frame.
Titles, credits, logos and your closing card all go on in post. Do not ask the model to render readable text.
Diagnose by asking which input failed, not which sentence.
The face drifted. Add an anchor, do not re-roll. Re-rolling changes the sample, not the constraint. The mechanism is covered properly in keeping a character consistent across shots.
A prop changed shape. Fix the reference image, not the prompt. If an object is described in the prompt and controlled by a reference, you have two control sources fighting.
The framing missed. Re-roll, and simplify the shot to one action. This one genuinely is a numbers game.
Text came out garbled. Remove text from the generation entirely and add it in post.
Then run a comparison pass by eye. Put every clip next to the reference set and check them as a set, not one at a time.
In the 152-take run, automated scoring approved a clip where a theatre scene had been replaced by a human head in a hanging basket. The metrics called it sharp, stable and well-lit. Software does not catch content errors. You do.
Building a narrative piece? The AI drama generator hub collects the dramatic shapes this workflow usually serves. For where it sits in a wider production, see our filmmakers page.
No. Seedance 2.5 generates clips of 4 to 30 seconds. A short film is assembled from many of those clips in an editor. Any tool promising an end-to-end film is describing an assembly pipeline, not a single generation.
In DomoAI Omni Reference, 4 to 30 seconds per generation, at 480P or 720P. MiniMax H3, MiniMax's model on the same surface, runs 4 to 15 seconds at 768P to 2K.
Plan for 6 to 15 shots, plus re-rolls. Re-rolls are normal and should be in your schedule from the start. Nobody should quote you a guaranteed total.
Use one identity reference in every generation, and give it only the identity job. The character does not persist between generations — each one re-anchors from your references. That is a different mechanism from memory, and planning around it is what makes it work.
480P or 720P. Choose your aspect ratio from Auto, 3:4, 4:3, 9:16, 16:9 or 1:1 before you start, and keep it identical across every shot in the film.
Several products describe a straight line: script, scenes, audio, export. The line is real. What those pages leave out is that continuity is not a checkbox inside it.
Compare surfaces, not models. On any surface, a multi-shot piece is many generations that each re-anchor from references. The assembly is yours.
A tool that hides that has not solved it. It has moved the disappointment later, to the point where shot 7 does not match shot 2.
We would rather tell you the shot count up front.
Plan the shots. Build one reference contract. Let every cut re-anchor instead of asking a single generation to carry a whole scene. Then assemble and grade once.
Open Omni Reference and generate shot 1.
Recent articles
© 2026 DOMOAI PTE. LTD.
DomoAI