Push and pull
Dolly-in, push-in, pull-back, zoom. Moving toward a subject concentrates attention and builds tension; moving away reveals context and releases it. The most useful pair in the vocabulary, and the easiest to overdo.
Motion Control
AI camera motion control
Direct camera moves in AI video — dolly-in, orbit, crane-up, whip pan, handheld shake — using natural language prompts. Cinematic camera control without motion-path software.
MP4 or MOV, 2–15s each, 15s total, up to 200MB. Requires R2 with a public domain — the model fetches the file over the network.
Model
Reference videos 0/3
MP4 or MOV, 2–15s each, 15s total, up to 200MB. Requires R2 with a public domain — the model fetches the file over the network.
0/2000
Outputs
Audio
With audio
Synchronized sound, voice and music
Your generated video appears here. Describe a shot and hit Generate.
This page is a prompt-based camera-direction workflow. Describe a dolly-in, orbit, crane-up, whip pan, handheld move, or locked-off shot while supplying a source video as reference. The provider uses that language as guidance; it does not transfer an exact motion path.
Camera language can make a prompt more specific: state the move, its speed, the subject action, and where the shot should begin or end. The generated result may still choose a different path or timing, so treat the prompt as direction rather than a tracking or motion-control file. Review the clip and refine the wording when the move drifts.
Three families cover most of what a shot needs. Each says something different to the viewer.
Dolly-in, push-in, pull-back, zoom. Moving toward a subject concentrates attention and builds tension; moving away reveals context and releases it. The most useful pair in the vocabulary, and the easiest to overdo.
Arc, orbit, whip pan, tilt. Lateral movement describes an object in three dimensions — which is why almost every product film in existence orbits. A whip pan does the opposite job: it hides a cut.
Crane up for scale, handheld for immediacy and unease, locked-off for formality, rack focus to move attention within a static frame. These set the register of a shot before anything happens in it.
Name one move per shot. A prompt that asks to dolly in, orbit, crane up and rack focus in eight seconds is asking for chaos, and chaos is what comes back. Pick the single move that does the storytelling job and describe its speed as well as its type — "slow push-in" and "fast push-in" are different shots with different emotional readings, and speed is the qualifier people most often leave out.
Then give the move something to work against. A camera orbiting an empty frame reads as a screensaver; a camera orbiting a subject that is itself doing something reads as filmmaking. Pair the move with a subject action and, where it matters, a starting and ending relationship — "begins wide on the doorway, pushes in to a close-up of the handle" tells the model both the move and its destination. Because audio is generated with the picture, a described camera move also shapes the sound: a fast whip pan brings its own whoosh.
"Slow orbit, left to right" is directable. "Dynamic cinematic camera movement" is not — it names no move and no speed, so the model chooses both.
Stating where the shot starts and where it ends up produces far more controlled results than naming the move alone.
Model specifications
Published limits, not marketing rounding. Every number below is enforced by the generator before your credits are reserved, so a job that would exceed a limit is rejected instead of failing halfway.
The model is ByteDance’s. What FrameAI adds is the part that decides whether it is usable at work: honest limits, honest billing, and output you are allowed to ship.
Text, up to nine reference images and up to three reference video clips are fused in one generation, not stitched afterwards. There is no separate text-to-video mode to switch into — you add whatever material you have and the model reconciles it.
Seedance 2.0 writes footsteps, room tone, impacts and score alongside the frames, locked to the action. Most models hand you a silent clip and leave the sound design to you; here the export is already finished.
One prompt covers a 9:16 vertical cut for Reels, a 16:9 master for YouTube and a 21:9 anamorphic frame for a title sequence. Resolution runs from 480p up to 4K, so the same take can serve a cinema-shaped frame or a phone.
The exact credit cost is calculated from your duration, resolution and ratio and shown before you press generate — never after. If a generation fails on the provider side, the reserved credits are returned automatically.
Seedance 2.0 Fast and Seedance 2.0 Mini exist for the twenty drafts nobody sees. Rough a shot at 480p for a fraction of the credits, then re-run the take that worked at full resolution on the flagship model.
Paid-plan output may be used commercially under the Terms of Service. Free-credit output is for evaluation only.
Same generator, one deliberate addition to the prompt.
Ask what the camera should make the viewer feel — pressure, scale, urgency, formality — and pick the move that does that job. One per shot.
Slow or fast, and from where to where. Then describe the subject action the move is built around, so the camera has something to be about.
If the move reads wrong, adjust the speed qualifier before rewriting anything else — it is usually the variable at fault.
In each of these the move is not a flourish — it is the reason the shot works.
A controlled orbit or push-in is the standard grammar of showing an object properly. Random drift makes the same product look incidental rather than considered.
Crane and pull-back moves are how landscapes, buildings and crowds read as large. Without the move, scale simply does not register.
A slow push-in during a held moment does what no amount of description can. Camera speed is the most direct control you have over how a clip feels.
It is a prompt-based workflow for describing camera movement. You name the move and its speed, optionally provide a source video, and the provider generates a new clip using that information as guidance.
You can describe common camera language such as dolly or push-in, pull-back, orbit, pan, whip pan, tilt, crane, handheld, locked-off, and rack focus. These are prompt cues, not a formal motion-path input.
You can, but it rarely helps. A four-to-fifteen-second clip has room for one move executed well. If a sequence needs several, generate them as separate shots and join them with Video Transition.
Say it — "slow push-in", "fast whip pan", "gentle orbit". Speed is the qualifier most often omitted and the one that most changes how a shot reads, so it is worth stating on every camera instruction.
Yes. This page requires a reference video, which gives the provider source motion and visual context; the camera words in your prompt add guidance. Neither input guarantees that the subject or path will remain identical.
The provider generates a new clip and may interpret the wording differently. Make the prompt concrete — for example, "slow dolly-in from wide to medium" — then regenerate and compare when the first result drifts.
Each tool is the same Seedance 2.0 model pointed at a different job. Move between them freely — your credits, history and exports are shared.
Write one line, pick a ratio, press generate. No editing suite, no render farm, no post-production pass — a finished clip with sound, ready to download.