Reference-to-video (R2V) is a generative AI technique that uses one or more reference images to guide video generation — not as the literal first frame, but as an identity anchor. The model studies your references (a character's face, a product's design, an art style) and generates new scenes where that identity stays consistent.
Reference vs. first frame
This is the key distinction from image-to-video:
- Image-to-video treats your image as the opening frame. The video starts exactly there.
- Reference-to-video treats your images as a definition of what things look like. The model can then place that character or product into an entirely new scene, angle, or action described by your prompt.
In short: I2V continues a picture; R2V casts it.
("R2V" is simply shorthand for reference-to-video — you'll also see it written as reference2video or ref-to-video.)
T2V vs I2V vs R2V at a glance
| Mode | Your input | The model treats it as | Best for |
|---|---|---|---|
| Text-to-video (T2V) | a written prompt only | the full creative brief | net-new scenes from scratch |
| Image-to-video (I2V) | one image (plus an optional last frame) | the literal opening frame | animating a still you already have |
| Reference-to-video (R2V) | a set of images, clips, and audio | identity and style anchors | consistent characters, products, and series |
Why it matters
Consistency is the hardest problem in AI video. Generate the same "red-haired girl in a yellow raincoat" twice from text alone and you'll get two different girls. Reference-to-video solves this:
- Character consistency — keep the same protagonist across shots, scenes, and episodes.
- Product fidelity — show your actual product from new angles, in new environments.
- Style continuity — carry an illustration style or brand look through a whole series.
Prompting tips
Give the model clean, well-lit references that show the subject clearly. Then let the prompt do the directing: new setting, new action, new camera. Say what should change — the references already say what should stay. More in our prompt writing guide.
Reference-to-video on Molyin
Reference-to-video is a standard mode across Molyin's video lineup rather than a single-model feature. Reference budgets per generation:
| Model | Reference budget | Notes |
|---|---|---|
| Seedance 2.5 / 2.0 | up to 9 images + 3 video clips + 3 audio tracks | audio references can drive voice or score |
| Wan 3.0 | up to 10 images + 5 videos + 5 audio | all-in-one reference with up to 30-second output |
| Wan 2.7 | up to 5 images and videos combined | no audio reference slot |
| MiniMax H3 | up to 9 images + 3 videos + 3 audio | reference clips capped at 15 seconds total |
| Kling 3.0 | 2–4 element images | named-element casting |
| FLUX 3 | 3–10 keyframe images | storyboard semantics — keyframes pin moments in time rather than identity |
Pick a model, drop in your references, and mention them inline ("@Image1 walks through the alley from @Video1") to direct the shot.
Related terms
- Text-to-Video — generate from a written prompt alone.
- Image-to-Video — animate an image as the literal first frame.
Try it with the AI video generator.