Reference to video AI: steer a render with your own images, clips and sound
Reference-to-video AI renders a new clip that is steered by media you attach: images that fix what a character, a product or a place looks like, clips that show a motion or a camera move, and audio that sets the sound. It sits between text-to-video, where the engine invents everything from your words, and image-to-video, where one picture becomes the first frame. References do not become the clip. They tell the engine what to keep. This guide explains what each kind of reference is good for, how many PrismPoster's engines accept, and how to write the prompt around them.
What is reference to video AI?
Reference to video is video generation guided by attached examples. You write a prompt for the new shot and attach images, clips or tracks as evidence of how things should look, move or sound; the engine renders new footage that matches them rather than copying them frame for frame.
The difference from image-to-video matters. In image-to-video, the uploaded picture is the start of the clip, so the composition is fixed. With references, the character from your photo can appear in a new place, from a new angle, in a new light. That is why references are the main tool for keeping a character or a product consistent across several shots.
How many references can a render use?
| Premium engine | Standard engine | |
|---|---|---|
| Reference images | Up to 30 | Up to 7 |
| Reference clips | Up to 10, 30 seconds combined | Up to 3 |
| Reference tracks (sound) | Up to 10 | Not supported |
| Start and end frames | Yes | Yes |
| Single take | 4 to 30 seconds | 4, 6, 8, 10 seconds |
Both engines are in the same Video studio. When you attach a reference track, the composer switches to Premium, because only Premium can take sound as a reference.
Which references should you attach?
Match the reference to the thing you want held steady:
- A character: three to five images of the same person or figure, different angles, neutral light. One face-on, one three-quarter, one full body. Consistency comes from agreement between images, so do not mix outfits unless the outfit should change.
- A product: the front, the back, the label close up and the product in a hand for scale. See our AI commercial generator guide for a full product spot.
- A place or a style: two or three frames of the location, or of the look you want (colour, lens, grain).
- A motion: a short clip of the move, such as a whip pan, a slow orbit, or a dancer's step. The engine reads the motion, not the content, when your prompt says so.
- A sound: a reference track for mood, tempo or ambience. Keep it short and on topic; a busy song steers less clearly than a single texture.
A real, identifiable person can only appear through PrismPoster's consented likeness setup. Photos of other people should not be attached as references.
How do you write the prompt around references?
Name what each reference is for. Engines follow references far better when the prompt tells them which reference carries which job:
The portrait references are the woman: keep her face and red coat.
The reference clip is the camera move: a slow orbit left.
0-6s: she stands on a rooftop at dusk, city lights behind her
6-14s: the camera orbits as she turns toward the skyline
14-20s: wind catches the coat; hold on her profile
Three habits help:
- Describe the new shot, not the reference. The engine can see the photo. Spend the words on what is different: the place, the action, the light.
- Keep references consistent with each other. Two photos of a product in different colours give the engine a choice you did not intend.
- Use start and end frames for exact compositions. A reference steers; a start or end frame fixes the image exactly. Our start and end frames guide covers when to use which.
When is reference-to-video the wrong tool?
- When you need one exact image to move. Use image-to-video; the AI image to video tool is built for that.
- When a shot has nothing to keep. A generic landscape or an abstract loop needs no references, and extras only constrain the engine.
- When the reference is someone else's work. Referencing a copyrighted clip or a real person's photos without permission is the same problem it would be in any edit.
What does a reference render cost?
References do not change the price. A Premium render bills per second of output: 18 credits a second at 480p, 40 at 720p, 50 at 1080p and 70 at 4K, so a ten-second take at 720p is 400 credits. Any credit pack opens Premium; before a first purchase, one 5-second draft at 480p is available for 90 credits. The full breakdown is on the AI video generator page.
Frequently Asked Questions
What is the difference between reference to video and image to video?
In image-to-video the picture becomes the first frame of the clip. In reference-to-video the images guide what things look like, and the engine composes a new shot around them.
How many reference images can I use?
Up to 30 on PrismPoster's Premium engine and up to 7 on the Standard engine, plus reference clips on both and reference tracks on Premium.
Can I use a video as a reference for motion?
Yes. Attach a short clip and say in the prompt that it is the motion or camera reference. On Premium, reference clips share 30 seconds between them.
Can I use music as a reference?
On the Premium engine, yes: up to 10 reference tracks steer the sound rendered with the picture.
Do references keep a character consistent across clips?
They are the most reliable way to. Attach the same set of character images to every shot in the sequence; our character consistency guide covers building that set.
Turn ideas into high-definition AI video clips & timeline cuts
Generate text-to-video and image-to-video in 360p up to 4K, extend scenes per second, and assemble complete multi-track videos with auto-captions and zero desktop install.