50% OFFClaim 50% Off
Nano Banana AINano Banana AI

Nano Banana 2.1 is live — sharper edits, cleaner text, up to 4K

Reference to Video

Multi-reference generation — anchor characters, wardrobe, and art direction, then generate shots that stay in the same world.

See pricing
Text to Image Image to Image Multi-Reference Editing Character Consistency Up to 4K
5 credits
0/2000

Avoid copyrighted characters, logos, music and sensitive or explicit content — it may be blocked.

Resolution
Aspect Ratio
1

Estimated cost: 5 credits

Your creation appears here

Pick a model, describe the shot, and hit Generate.

Why Reference to Video

Why creators choose Reference to Video

Reference to video is for work that needs more than a single clip: recurring characters and a steady look.

Identity that holds

Faces, outfits, and palette carry across shots.

Many references

Up to 30 images on Seedance 2.5, nine on MiniMax H3 Max.

Video and audio cues

Seedance 2.5, Wan 3.0, and MiniMax H3 Max accept clip and audio references.

Reusable sets

The same references work across every shot.

Reference to video guide

How reference to video keeps characters consistent

Reference to video generates clips from several reference images, and on some models from video and audio cues as well. It exists to solve one problem: keeping the same character, outfit, and look from one shot to the next. This page explains how to use it.

How reference to video works

Instead of a single first frame, you provide a set of references that define the subject and style, plus a prompt for the new shot. Seedance 2.5 accepts up to 30 reference images, 10 reference videos, and 10 audio clips. MiniMax H3 Max accepts up to nine images, three videos, and three audio clips. Wan 3.0 takes up to 10 images, 5 videos, and 5 audio clips.

Building a reference set

Choose references that agree with each other. Use one clear face, the outfit, and the setting, in similar lighting. If a reference contradicts another, the model has to compromise and consistency suffers. In the prompt, say what each image contributes, such as "the jacket from reference two".

Character

A clear view of the face and body.

Wardrobe and props

The items that must stay the same.

Setting and style

The place, palette, and art direction.

Making a sequence

Reuse the same reference set for every shot in a scene and change only the prompt. Generate an establishing shot, a medium shot, and a close-up with the same references, and the sequence will read as one world. Draft at low resolution first, then re-run the shots you keep.

How to use Reference to Video

How to use Reference to Video in three steps

Pick agreeing images for character, wardrobe, and setting.

00:0000:1500:30

Who uses Reference to Video

Who uses Reference to Video for image work

Short-film makers

Short-film makers

Scenes with the same cast throughout.

Brand storytellers

Brand storytellers

A consistent spokesperson across a campaign.

Fashion brands

Fashion brands

The same model and outfit across angles.

Series creators

Series creators

Episodes that keep a recognizable look.

FAQ

Reference to Video FAQ

What is reference-to-video?
Instead of a single first frame, you supply multiple reference stills (and for some models, video or audio cues) that define the subject, style, and mood of the generated clip.
How many references can I add?
Seedance 2.5 accepts up to 30 reference images per shot; MiniMax H3 Max takes up to 9.
Does it keep characters consistent?
That is its core purpose — identity, outfit, and palette carry across every generated clip.
Can I mix reference types?
Reference mode accepts still images; Muse Video is stills-first.

Your next visual is one prompt away

Free credits on signup. No card required.