Reference to Video

Reference-guided video AI

Attach up to 9 reference images to lock character appearance, product details, or art direction, then describe the shot. Get consistent AI video output guided by your real assets.

Use up to 9 reference images to guide a character, product, or look; generated details may still vary.

Model

Input modeReferences

Reference images 0/9

Reference videos 0/3

MP4 or MOV, 2–15s each, 15s total, up to 200MB. Requires R2 with a public domain — the model fetches the file over the network.

0/2000

Aspect ratio
5
4s15s
Resolution

Outputs

Audio

With audio

Synchronized sound, voice and music

Credits required:36

Your generated video appears here. Describe a shot and hit Generate.

My creationsDownload

Generation tips

  • - Attach up to 9 reference images to lock character appearance, product details, or art direction, then describe the shot. Get consistent AI video output guided by your real assets.
  • - Use up to 9 reference images to guide a character, product, or look; generated details may still vary.
  • - Generation time varies with the model, clip settings, and provider load.
  • - Use up to 9 reference images to guide a character, product, or look; generated details may still vary.
  • - Name the subject, the action, and the camera move. Vague prompts give vague shots.
  • - Longer clips cost more credits — start at 5s to test a look, then extend.
  • - Audio is generated with the picture; describe the sound you want in the prompt.

Reference to video — guide a subject, product, or visual direction

Reference to video adds source material to a generation request. Attach up to nine images and three video clips, describe the shot, and let the provider use those assets as visual guidance. The result is newly generated, so identity, lighting, motion, and fine detail may still change.

Use references when a prompt alone does not describe the subject or look well enough. A character sheet, product image, or palette board can give the provider more context, but it is not a post-production lock or a guarantee of consistency. Keep the set focused, describe the shot separately, and review the generated take before using it in a sequence.

  • Up to nine images and three video clips
  • References used as generation guidance
  • Review identity and details in every result

Three things worth locking down

A reference is an instruction to hold something constant. These are the three that most often decide whether footage is usable.

A character who stays one person

Supply two or three clear images of the same face — different angles, consistent lighting — and the model keeps that identity across the clip and across separate generations. This is what makes a series of shots feel like one production.

A product that stays your product

Colourway, proportions, logo placement, materials. Text prompts invent a plausible object; a reference image pins the real one, which is non-negotiable for anything commercial.

An art direction that holds

Palette, grain, contrast, era, lens character. Feeding in stills that share a look transfers that look, so a batch of clips reads as a coherent campaign instead of a mood-board accident.

How to choose references that actually work

The instinct is to upload everything you have. Resist it — references compete with each other, and nine conflicting images produce a worse result than three agreeing ones. Pick images that share lighting logic and are unambiguous about the thing you want held: for a face, two or three angles of the same person under similar light; for a product, a clean pack shot plus one detail; for a look, stills that genuinely share a grade rather than merely a vibe. Crop out anything you do not want carried, because the model has no way of knowing that the busy background was incidental.

Video references work differently from image references, and it is worth being deliberate about which you reach for. An image says *what things look like*; a video clip says *how things move* — pacing, camera behaviour, the rhythm of an action. Attaching a clip whose motion you admire lets you apply that movement to a scene the clip never contained. Combined budgets apply: up to three reference videos totalling no more than fifteen seconds, so choose the segment that carries the move rather than uploading a whole take.

Three strong references beat nine weak ones

Conflicting lighting or contradictory angles force the model to average. Consistency between your references is what produces consistency in the output.

Images carry look, videos carry motion

Use stills to fix identity and art direction; use clips to borrow a camera move or an action’s rhythm. Mixing both in one request is where this tool is strongest.

Model specifications

What Seedance 2.0 actually delivers on FrameAI

Published limits, not marketing rounding. Every number below is enforced by the generator before your credits are reserved, so a job that would exceed a limit is rejected instead of failing halfway.

Up to 4K
Output resolutions — 480p / 720p / 1080p / 4K
4–15 sec
Clip length per generation, at 24 fps
Native audio
Sound effects and ambience generated in sync
9 + 3
Reference images plus reference videos, one request
7 ratios
Supported aspect ratios — Auto, 16:9, 4:3, 1:1, 3:4, 9:16, 21:9
Commercial use on paid plans
Subject to the Terms of Service

Why creators run Seedance 2.0 on FrameAI

The model is ByteDance’s. What FrameAI adds is the part that decides whether it is usable at work: honest limits, honest billing, and output you are allowed to ship.

Every input in a single pass

Text, up to nine reference images and up to three reference video clips are fused in one generation, not stitched afterwards. There is no separate text-to-video mode to switch into — you add whatever material you have and the model reconciles it.

Sound generated with the picture

Seedance 2.0 writes footsteps, room tone, impacts and score alongside the frames, locked to the action. Most models hand you a silent clip and leave the sound design to you; here the export is already finished.

Seven ratios, up to 4K

One prompt covers a 9:16 vertical cut for Reels, a 16:9 master for YouTube and a 21:9 anamorphic frame for a title sequence. Resolution runs from 480p up to 4K, so the same take can serve a cinema-shaped frame or a phone.

Credits you can predict

The exact credit cost is calculated from your duration, resolution and ratio and shown before you press generate — never after. If a generation fails on the provider side, the reserved credits are returned automatically.

Built for iteration, not for queuing

Seedance 2.0 Fast and Seedance 2.0 Mini exist for the twenty drafts nobody sees. Rough a shot at 480p for a fraction of the credits, then re-run the take that worked at full resolution on the flagship model.

Yours to sell

Paid-plan output may be used commercially under the Terms of Service. Free-credit output is for evaluation only.

Generate a reference-driven shot in three steps

Set up the references once and every subsequent shot in the series inherits them.

  1. 1

    Attach your references

    Upload up to nine images and three clips. Crop each one to the thing you want held, and drop anything that fights the others on lighting or angle.

  2. 2

    Describe only the new part

    The references already say what things look like. Your prompt should say what happens and how the camera behaves — re-describing the reference wastes the instruction.

  3. 3

    Generate the series

    Keep the same reference set and change only the prompt to produce a second, third and fourth shot that all belong to the same world.

Work that only becomes possible with references

Each of these fails immediately on a text-only model, for the same reason: something has to stay exactly itself.

Brand campaigns with a recurring face

A spokesperson or mascot appearing across a dozen clips has to be the same person every time. A locked character reference makes a campaign out of what would otherwise be a dozen strangers.

Catalogue video at scale

Hundreds of SKUs, each needing motion, each needing to be accurate. Swap the product reference, keep the prompt and the art direction, and the whole catalogue matches.

Episodic and serialised content

Recurring characters and locations across episodes need continuity that a prompt cannot guarantee. References give a series its through-line.

Reference to video — frequently asked questions

What is reference to video?

It is video generation that accepts source images or videos as visual guidance. You can attach up to nine reference images and three reference video clips, then describe the shot you want to generate.

How many references can I use at once?

Up to nine images and up to three video clips in a single request, with the combined reference video length capped at fifteen seconds. More is not automatically better — references that disagree with each other degrade the result.

How consistent is character identity really?

There is no fixed consistency guarantee. Two or three clear references of the same subject under compatible lighting can provide better guidance, while extreme angles, occlusion, or conflicting references can increase drift.

What is the difference between image and video references?

Image references fix appearance — identity, colour, materials, art direction. Video references contribute motion — camera behaviour, pacing, the rhythm of an action. You can use both in the same request, and doing so is where the tool is strongest.

Can I reuse the same references across generations?

Yes. Keeping the reference set fixed and changing only the shot prompt is a sensible way to guide a related set of clips, but inspect each result because separate generations can still differ.

Can I use photos of real people as references?

Only with that person’s consent. The model has no way to verify permission, so the responsibility sits with you — and generating identifiable people without consent is not covered by your commercial licence.

Keep going with the rest of the toolkit

Each tool is the same Seedance 2.0 model pointed at a different job. Move between them freely — your credits, history and exports are shared.

Generate your first Seedance 2.0 video now

Write one line, pick a ratio, press generate. No editing suite, no render farm, no post-production pass — a finished clip with sound, ready to download.

  • Free credits on sign-up
  • Paid commercial use
  • Refunded if a generation fails