
Table of Content

Try DomoAI, the Best AI Animation Generator
Turn any text, image, or video into anime, realistic, or artistic videos. Over 30 unique styles available.
Character consistency comes from role-assigned reference images, not from repeating a description in the prompt. In DomoAI Omni Reference, you give each reference exactly one job, and every generation re-anchors from that set. Longer prompts do not fix drift. Anchors do.
Seedance 2.5 is ByteDance's model, integrated by DomoAI in Omni Reference. What follows is how its reference slots actually behave.
Each generation is a fresh sample. Nothing carries over from the last one unless you hand it back in.
So without a re-anchor, small differences accumulate. The jaw narrows a little. The hair sits slightly differently. By clip four you are looking at a cousin.
Creators outside DomoAI describe face, voice and wardrobe drift as the single most common problem in multi-shot AI video. The clearest statement of the fix comes from a thread on r/generativeAI that tells people to stop prompting for it: pure text-to-video continuity is described there as "rolling the dice."
That is the hinge of this page. The fix is an anchor image, not a better sentence.
Three tools, one order, and the order matters.
Midjourney generates the identity anchor. One clean image, front-facing, full body, neutral studio background, no readable text or logos.
GPT Image 2 generates variants from that anchor — different wardrobe, different location, different light, same person.
Seedance 2.5 generates the video from those images.
Never generate the first identity image in GPT Image 2. And generate every variant from the anchor, never from the last variant you made. Deriving variant C from variant B stacks two rounds of drift onto the same face before the video model has touched it.
Midjourney — identity anchor
an original adult character in their late twenties with a distinctive asymmetric
bob and calm, level expression, full-body front-facing identity portrait, plain
olive knit top and dark straight-leg trousers, plain light-grey studio backdrop,
even softbox lighting, realistic skin and fabric texture, clean character
reference, generous negative space --ar 9:16 --style raw --s 70
--no text, logo, celebrity, copyrighted character, extra person, watermarkGPT Image 2 — location variant, generated from the anchor
Use the supplied portrait as the exact identity reference. Create a vertical
full-body keyframe of the same person in a rain-wet night market street. Preserve
the exact face, age, hairstyle, skin tone, body proportions and recognizable
identity from the reference image. Use practical light from overhead market
stalls, realistic skin and fabric texture, and enough space around the subject
for animation. Change only the environment and matching light. No text, logos,
brand names, extra people or watermarks. Portrait orientation, 9:16.This is the mechanic the model-agnostic tutorials skip.
| Slot | Its one job | What it must not do |
|---|---|---|
Image 1 | identity, and the literal first frame, and the location and light direction | — |
Image 2 | facial identity, hairstyle, wardrobe | donate framing, composition or lighting |
Image 3 | a prop, or a space | donate identity |
Video 1 | movement timing and camera rhythm | donate identity or wardrobe |
Audio 1 | soundtrack or dialogue timing | drive cut timing |
References are addressed by upload order. Image 1 is the first image you uploaded, Image 2 the second, and so on.
Two consequences worth writing on a sticky note. Image 1 is the opening frame, so it also fixes where the camera starts and where the light comes from. And any faces-only reference needs an explicit exclusion clause, or it will fight the first frame for the composition.
Seedance 2.5 — the reference contract paragraph
Image 1 controls the location, camera position, light direction and composition,
and is the literal first frame.
Image 2 controls the character's facial identity, hairstyle and wardrobe ONLY.
Do not recreate its framing, composition, or lighting.
Image 3 controls the exact prop only. Take no identity or lighting from it.Use the minimum sufficient set. Omni Reference accepts 30 images, 10 videos and 10 audio files. Every reference beyond the ones with a named role adds ambiguity, not control.
For a person, lock five things: face, hair, body scale, outfit, speaker identity.
For an object, the list is different: silhouette, closure, material boundaries, label area, proportions, colour. That distinction matters more than it sounds. A car is an object, which is why its consistency problem is a product problem and gets solved with the second list.
Write the locked attributes once, in one place. If a jacket is controlled by Image 2, do not also describe the jacket in the prompt. Two descriptions are two control sources, and that is where wardrobe drift starts.
The strongest evidence on this page comes from a 152-take same-prompt comparison run in ComfyUI between LTX-2.5 and MiniMax H3. The tester's finding:
"Identity drift is a single-constraint problem. With only a first-frame anchor, the host's face slides toward a generic one mid-clip. New seeds do NOT fix it. Double anchor (FLF2V) does: worst clips went from −0.53 to −0.10 vs reference."
Those are other models on another surface. Do not read it as a Seedance 2.5 benchmark. Read it as the shape of the problem, which is the same everywhere.
Two things follow, and both are actionable here.
Re-rolling is not a fix. It is the first thing most people try, and it changes the sample without changing the constraint.
A second anchor point is. Omni Reference supports a first frame and a last frame. That is exactly the constraint that test was adding.
The failure mode at the other extreme is worth knowing too. On a months-long retro project, one creator abandoned motion transfer entirely: "The body movement worked surprisingly well but the faces did not. They would gradually morph until the actors stopped looking like themselves." Only one or two shots from that workflow survived.
Run a static shot, a motion shot, and a location change. In that order.
If the face holds across those three, it will hold across nine. If it does not, the anchor is the problem, and no amount of prompt editing on shots four through nine will rescue it.
This is fifteen minutes that routinely saves an afternoon.
Same person, four separate generations, four different places. The identity slot never changes; the location slot changes every time.
| Shot | Image 1 | Image 2 | Duration |
|---|---|---|---|
| 1 | studio, front light | identity anchor | 6s |
| 2 | night market street | identity anchor | 8s |
| 3 | stairwell, overhead practical | identity anchor | 6s |
| 4 | rooftop, low sun | identity anchor | 10s |
Seedance 2.5 — Omni Reference — shot 2
Image 1 controls the night market street, camera position, light direction and
composition, and is the literal first frame.
Image 2 controls the character's facial identity, hairstyle and wardrobe ONLY.
Do not recreate its framing, composition, or lighting.
Create an 8-second vertical 9:16 shot. The character walks between two stalls and
stops to look at something off-frame left. The camera tracks with them at chest
height and settles when they stop. Overhead practical light from the stall
canopies, warm on the face, hard shadow under the brow. Restrained grain,
realistic skin texture with visible pores as the camera settles.
Preserve the exact face, age, hairstyle, skin tone and body proportions from
Image 2. Exactly one person appears in this video. No readable text, logos,
signage or subtitles.Four failures and their actual causes.
The face drifted mid-clip. Add a second anchor point. Do not re-roll.
The reference hijacked the composition. Your faces-only image is missing the exclusion clause. Add "do not recreate its framing, composition, or lighting."
The wardrobe changed between shots. Something is described in two places at once. Find the object that appears both in a reference and in the prompt, and delete one.
Text came out garbled on clothing or signage. Remove text from the generation and add it in post.
One recovery move worth learning: a half-working generation is still a reference source. Grab a single frame from it, clean it up, and use that frame as the starting image for a different shot. The same creator quoted above does this routinely, in both directions — including pulling a frame from a failed video to use as an end frame.
But do not chain last frames repeatedly. Quality degrades fast that way, and the same creator describes it as the classic repeated-recompression problem, fried by the fourth iteration. The fix is a scene transition or a camera-angle change, because that lets you start from a fresh frame. Which is the same mechanism as assembling a short film from separate generations: a cut re-anchors.
Use keyframes when you already know what the frames look like. Use references when you already know who the character is.
Frames to Video (2–8 keyframes) — you have specific images and you want the video to pass through them in order. The frames define the states; the prompt defines the transitions between them. Best when the shot is one continuous action with a known start, middle and end.
Omni Reference (role-assigned reference images) — you have an identity and you want it to survive across separate generations that are not a single continuous action. The references define who and where; each generation re-anchors from them. Best when the job is many shots, many locations, or a cut between them.
They are not competitors. A film that needs both uses keyframes inside a shot and references across shots.
So when you already have the frames, use the keyframe method for a music-video reel. The mechanic itself is covered in turn multiple keyframes into one video.
Give one reference image the identity job and include it in every generation. Add the exclusion clause so it does not also control composition. The character does not persist between generations — each one re-anchors from your references.
Up to 30 images, 10 videos and 10 audio files. Use far fewer. Every reference without a named role adds ambiguity rather than control.
Yes. Image 1 is the literal first frame, so it also fixes the location, camera position and light direction. Other references contribute only the job you assign them.
References when you know who the character is and need them across separate shots. Keyframes when you already have specific images and want one continuous action to pass through them.
Usually because the outfit is described in the prompt and controlled by a reference at the same time. Two control sources compete. Describe it in one place only.
Yes. Omni Reference accepts JPG, JPEG, PNG, WEBP, HEIC and HEIF files from any source. A Midjourney anchor is the recommended starting point in this workflow.
Character sheets and LoRA training both solve consistency by front-loading effort into a model or an asset pack. They work, and they cost time before you generate anything.
The reference method front-loads less. One anchor, a couple of variants, and a contract paragraph you reuse. It is suited to projects where the character is settled but the shot list is still moving.
LoRA training is not a DomoAI feature — it is what some other workflows do. If your character is going to appear in hundreds of shots over months, that tradeoff may be worth it elsewhere. For a music video, a scene or a short film, anchors are usually enough.
If you want the same treatment for a different model, we wrote MiniMax H3 character consistency as its sibling. For motion-led work driven by an existing performance, look at Character to Video instead.
Build one anchor. Give every reference one job. Add the exclusion clause. Then change the location, the wardrobe or the action — one at a time, and never the identity slot.
Open Omni Reference and test three clips.
Recent articles
© 2026 DOMOAI PTE. LTD.
DomoAI