Reference-to-Video (R2V)

July 11, 2026
Reference-to-video (R2V) uses reference images to keep characters, products, and styles consistent in AI video — and how it differs from image-to-video.
AI Video

Reference-to-video (R2V) is a generative AI technique that uses one or more reference images to guide video generation — not as the literal first frame, but as an identity anchor. The model studies your references (a character's face, a product's design, an art style) and generates new scenes where that identity stays consistent.

Reference vs. first frame

This is the key distinction from image-to-video:

  • Image-to-video treats your image as the opening frame. The video starts exactly there.
  • Reference-to-video treats your images as a definition of what things look like. The model can then place that character or product into an entirely new scene, angle, or action described by your prompt.

In short: I2V continues a picture; R2V casts it.

("R2V" is simply shorthand for reference-to-video — you'll also see it written as reference2video or ref-to-video.)

T2V vs I2V vs R2V at a glance

ModeYour inputThe model treats it asBest for
Text-to-video (T2V)a written prompt onlythe full creative briefnet-new scenes from scratch
Image-to-video (I2V)one image (plus an optional last frame)the literal opening frameanimating a still you already have
Reference-to-video (R2V)a set of images, clips, and audioidentity and style anchorsconsistent characters, products, and series

Why it matters

Consistency is the hardest problem in AI video. Generate the same "red-haired girl in a yellow raincoat" twice from text alone and you'll get two different girls. Reference-to-video solves this:

  • Character consistency — keep the same protagonist across shots, scenes, and episodes.
  • Product fidelity — show your actual product from new angles, in new environments.
  • Style continuity — carry an illustration style or brand look through a whole series.

Prompting tips

Give the model clean, well-lit references that show the subject clearly. Then let the prompt do the directing: new setting, new action, new camera. Say what should change — the references already say what should stay. More in our prompt writing guide.

Reference-to-video on Molyin

Reference-to-video is a standard mode across Molyin's video lineup rather than a single-model feature. Reference budgets per generation:

ModelReference budgetNotes
Seedance 2.5 / 2.0up to 9 images + 3 video clips + 3 audio tracksaudio references can drive voice or score
Wan 3.0up to 10 images + 5 videos + 5 audioall-in-one reference with up to 30-second output
Wan 2.7up to 5 images and videos combinedno audio reference slot
MiniMax H3up to 9 images + 3 videos + 3 audioreference clips capped at 15 seconds total
Kling 3.02–4 element imagesnamed-element casting
FLUX 33–10 keyframe imagesstoryboard semantics — keyframes pin moments in time rather than identity

Pick a model, drop in your references, and mention them inline ("@Image1 walks through the alley from @Video1") to direct the shot.

Try it with the AI video generator.

Reference-to-Video (R2V) | Molyin