Character identity
Keep faces, hair, body proportions, and distinctive features recognizable across new scenes.
Reference to Video AI
Upload reference images, videos, or audio and generate smooth, coherent videos where your character speaks and acts with perfect facial consistency and natural motion. No more reshoots!
Choose a model, add compatible image, video, or audio references, then mention them in your prompt.
Reference assets
9 images · 3 videos · 3 audio files
Choose a template to load its prompt and reference direction.
08 reference to video use cases
See how reference to video keeps identity, wardrobe, products, and visual worlds recognizable across drama, UGC, anime, fashion, music, and game scenes.
Explore fashion, product, performer, and mascot continuity in the original eight reference to video cases.
Reference-first generation
Reference to video is an AI generation workflow that uses one or more source assets as creative evidence. Instead of asking a text prompt to remember a face, product, outfit, movement, camera style, and sound at once, reference to video lets each source define a specific part of the new shot.
A strong reference to video result does not reproduce a source frame for frame. It preserves the details that create continuity while allowing the setting, action, framing, and story beat to change. This makes reference to video useful for episodic creators, AI influencers, product campaigns, anime characters, fashion sequences, music videos, and branded worlds.
The difference is control: references stay active throughout the shot, defining what must remain recognizable and what may change.
Each source should own one production decision. Clear roles make the result easier to direct, review, and refine.
Keep faces, hair, body proportions, and distinctive features recognizable across new scenes.
Preserve product shape, color, materials, packaging, and label placement in new shots.
Use a short video to guide gestures, timing, body movement, and performance energy.
Reference a pan, orbit, push-in, handheld move, or other camera behavior that is hard to describe.
Carry a location, lighting language, palette, or visual world into a newly composed shot.
Guide dialogue, voice, music, beats, and action timing on models that accept audio references.
Start with the smallest useful set of sources. These two workflows separate appearance, motion, and audio so every reference has one clear job.
Use the image to anchor a person, product, outfit, or visual style. Use the video to guide movement, performance timing, framing, and camera direction. The planned example will be delivered at 720P.
@Image1 defines the subject and appearance. Follow the action and camera movement from @Video1 while creating a new scene.
Add audio when voice, music, rhythm, or synchronization is part of the brief. The image owns appearance, the video owns performance, and the audio owns sound and timing. The planned example will use H3's 768P output.
Keep the subject from @Image1, follow the performance in @Video1, and synchronize the reveal to @Audio1.
Reference to video consistency failures often come from conflicting sources, not a lack of prompt adjectives. Simplify the evidence, state what must stay fixed, and test one creative change at a time.
For reference to video identity control, use one clean front or three-quarter image, remove conflicting portraits, and describe the person once. Add a second angle only when it contributes new information.
Assign each person a separate reference token and repeat the mapping in the action sentence. Keep wardrobe descriptions attached to the correct name.
Give reference to video a clean product hero image plus an angle showing geometry, branding, and material. Ask it to preserve shape, label placement, color, and scale.
Trim source media to the exact movement or beat. In a reference to video prompt, state one job for every source and remove redundant assets.
Use a full-body reference with the complete outfit visible. Name the garment, material, color, and accessories as fixed details, and avoid conflicting wardrobe images.
Use a short video showing the exact pan, orbit, push-in, or handheld movement. Assign that clip only to camera behavior instead of asking it to define the subject as well.
Choose by what must remain under control. Text creates from a description, image to video animates an existing frame, and reference to video uses several sources as constraints for a newly composed shot.
Text to videoA written description
Image to videoOne image treated as the opening frame
Reference to videoSeveral assets with separate creative jobs
Text to videoConcept, subject, action, and mood
Image to videoComposition and appearance near the first frame
Reference to videoIdentity, product, motion, camera, style, and audio
Text to videoOpen-ended ideation and new scenes
Image to videoA single short animation from one still image
Reference to videoStories, campaigns, UGC series, character IP, and performance-led shots
Text to videoDescribe everything the model must create
Image to videoDescribe movement away from the uploaded frame
Reference to videoAssign roles to @Image, @Video, and @Audio sources
Build a controlled reference to video brief by assigning every source a specific production job before generating a new scene.
A reference to video project starts with clear images for identity, products, outfits, or style. Add short video and audio references when they contribute motion, camera, voice, or timing.
Tell the model what @Image1, @Video1, and @Audio1 control, then describe the action, setting, camera, and details that must not change.
Review the reference to video result for faces, wardrobe, product geometry, roles, motion, and audio timing. Change one instruction at a time.
The best reference to video model depends on your sources and the detail that matters most. The composer validates supported asset types and limits before it sends a job.
The complete Seedance lineup
Choose Seedance 2.0 Mini, 2.0 Fast, or 2.5. Mini supports 480P/720P, Fast supports 720P, and 2.5 supports 480P/720P with generation up to 30 seconds.
High-resolution multimodal reference
MiniMax H3 supports 768P and 2K with image, direct-video, and audio references. Images and audio add no extra charge; direct videos are billed by total input plus output duration.
Fast Gemini video generation
Omni Flash 1.1 supports 720P, 1080P, and 4K output at 8 or 10 seconds. It accepts up to seven weighted reference slots, with one video using two slots.
Prepare identity references, name the details that cannot change, and review face, outfit, and proportions across every shot.
Read the character guideGive appearance to images, movement to video, timing to audio, and connect every source with explicit reference tokens.
Read the multimodal guideReview identity consistency, scene continuity, image to video, multimodal prompting, and related production terms.
Open the glossary$144 billed yearly
Up to 37 videos
Credits issued monthly. No rollover.
$348 billed yearly
Up to 100 videos
Credits issued monthly. No rollover.
$1,068 billed yearly
Up to 325 videos
Credits issued monthly. No rollover.
No video reference
| Model / Resolution | Lite1,500 credits | Pro4,000 credits | Ultra13,000 credits |
|---|---|---|---|
| Seedance 2.0 Mini · 480P40 credits / 5s | Up to 37 videos | Up to 100 videos | Up to 325 videos |
| Seedance 2.0 Mini · 720P80 credits / 5s | Up to 18 videos | Up to 50 videos | Up to 162 videos |
| Seedance 2.0 Fast · 720P180 credits / 5s | Up to 8 videos | Up to 22 videos | Up to 72 videos |
| Seedance 2.5 · 480P190 credits / 5s | Up to 7 videos | Up to 21 videos | Up to 68 videos |
| Seedance 2.5 · 720P380 credits / 5s | Up to 3 videos | Up to 10 videos | Up to 34 videos |
| MiniMax H3 · 768P110 credits / 5s | Up to 13 videos | Up to 36 videos | Up to 118 videos |
| MiniMax H3 · 2K160 credits / 5s | Up to 9 videos | Up to 25 videos | Up to 81 videos |
| Omni Flash 1.1 · 720P120 credits / 8s | Up to 12 videos | Up to 33 videos | Up to 108 videos |
| Omni Flash 1.1 · 1080P180 credits / 8s | Up to 8 videos | Up to 22 videos | Up to 72 videos |
| Omni Flash 1.1 · 4K350 credits / 8s | Up to 4 videos | Up to 11 videos | Up to 37 videos |
Total credits = base generation credits + video-reference credits
Without a native video reference
Base credits = output rate × output seconds
Text prompts, reference images, and reference audio are charged only for the generated duration.
With native video references
Total credits = base generation credits + video-reference credits
Seedance and H3 charge video references per second. Reference durations are added together before rounding up; images and audio are free. Omni uses a fixed video-reference add-on, not a per-second fee.
| Model | Resolution | 5s output · no video reference | 5s reference + 5s output |
|---|---|---|---|
| Seedance 2.0 Mini | 480P | 40 | 40 + 5 × 2 = 50 |
| Seedance 2.0 Mini | 720P | 80 | 80 + 5 × 4 = 100 |
| Seedance 2.0 Fast | 720P | 180 | 180 + 5 × 8 = 220 |
| Seedance 2.5 | 480P | 190 | 190 + 5 × 10 = 240 |
| Seedance 2.5 | 720P | 380 | 380 + 5 × 20 = 480 |
| MiniMax H3 | 768P | 110 | 110 + 5 × 22 = 220 |
| MiniMax H3 | 2K | 160 | 160 + 5 × 32 = 320 |
| Resolution | 8s output · no video reference | 3s reference + 8s output |
|---|---|---|
| 720P | 120 | 120 + 50 = 170 |
| 1080P | 180 | 180 + 60 = 240 |
| 4K | 350 | 350 + 70 = 420 |
Seedance and H3 charge video references per second. Reference durations are added together before rounding up; images and audio are free. Omni uses a fixed video-reference add-on, not a per-second fee.
Reference to Video
Use images for appearance, video for motion, and audio for sound and timing. Give every source a clear role, then direct what the next shot should change.