Subject and action
Who or what is on screen, and the single thing it does. One clear action per generation — a hand lifting a cup, a car pulling away — reads far better than a chain of events squeezed into a few seconds.
Text to Video
Text to video
Describe any scene — subject, action, camera move — and Seedance 2.0 renders it at up to 4K with matching audio. No editing software, no render farm, no plugins needed. Free to try.
Describe the scene, the character action, and the camera move. e.g. “slow dolly-in on a lone figure crossing a neon-lit street in the rain”.
Model
0/2000
Outputs
Audio
With audio
Synchronized sound, voice and music
Your generated video appears here. Describe a shot and hit Generate.
Text to video is the purest form of the tool: no upload, no source material, nothing but a description. Seedance 2.0 reads your sentence and produces a 4-to-15-second clip at up to 4K, with lighting, camera movement and physically plausible motion inferred from the words — plus audio generated in sync with what it decided to show you.
It works because the model was trained on how shots are actually constructed, not just on what objects look like. Ask for a handheld follow behind a cyclist at golden hour and you get camera shake, a trailing frame, long shadows and warm key light — three craft decisions you never spelled out. Prompts are understood in both English and Chinese, and paid-plan output may be used commercially under the Terms of Service.
Four levers, all of them reachable from plain language. Learning to pull them deliberately is the whole skill.
Who or what is on screen, and the single thing it does. One clear action per generation — a hand lifting a cup, a car pulling away — reads far better than a chain of events squeezed into a few seconds.
Dolly, pan, orbit, crane, locked-off; wide, 35mm, macro; handheld or on sticks. These terms are understood as instructions, and naming one is the fastest way to stop the model from choosing for you.
Golden hour, overcast, hard key, neon spill, candlelit. Light does more for how a clip feels than any adjective about beauty, and it is the lever most people forget to touch.
The most common failure is writing a story instead of a shot. A model with a fifteen-second ceiling cannot deliver "a woman wakes up, makes coffee, drives to work and meets her friend" — it will attempt all four and land none. Pick the single moment that carries the idea and describe it the way a shot list would: subject, action, camera, light, mood. If the sequence genuinely needs multiple beats, generate them as separate takes and join them with Video Transition.
The second lesson is that specificity beats enthusiasm. "Cinematic, stunning, 8K, masterpiece" tells the model nothing it can act on; "low-angle 35mm, wet asphalt reflecting neon, slow push-in, light rain" tells it five things. Then iterate one variable at a time — change the lens, keep everything else, and compare. That discipline turns prompting from guesswork into something you get reliably better at, and it is much cheaper to practise on Seedance 2.0 Mini at 480p before committing a take at 4K.
"Slow push-in on a chef plating scallops, tungsten key light from the left, shallow depth of field, steam rising" — subject, action, camera, light, lens, detail. Six decisions, one sentence.
Because audio is generated with the picture, naming what you want to hear — "sizzle, distant kitchen chatter, no music" — shapes the mix as well as the frame.
Model specifications
Published limits, not marketing rounding. Every number below is enforced by the generator before your credits are reserved, so a job that would exceed a limit is rejected instead of failing halfway.
The model is ByteDance’s. What FrameAI adds is the part that decides whether it is usable at work: honest limits, honest billing, and output you are allowed to ship.
Text, up to nine reference images and up to three reference video clips are fused in one generation, not stitched afterwards. There is no separate text-to-video mode to switch into — you add whatever material you have and the model reconciles it.
Seedance 2.0 writes footsteps, room tone, impacts and score alongside the frames, locked to the action. Most models hand you a silent clip and leave the sound design to you; here the export is already finished.
One prompt covers a 9:16 vertical cut for Reels, a 16:9 master for YouTube and a 21:9 anamorphic frame for a title sequence. Resolution runs from 480p up to 4K, so the same take can serve a cinema-shaped frame or a phone.
The exact credit cost is calculated from your duration, resolution and ratio and shown before you press generate — never after. If a generation fails on the provider side, the reserved credits are returned automatically.
Seedance 2.0 Fast and Seedance 2.0 Mini exist for the twenty drafts nobody sees. Rough a shot at 480p for a fraction of the credits, then re-run the take that worked at full resolution on the flagship model.
Paid-plan output may be used commercially under the Terms of Service. Free-credit output is for evaluation only.
The whole loop takes minutes, which is what makes iteration practical.
Name the subject, the single action, the camera move and the light. Skip adjectives that mean "good" — they carry no information the model can use.
Vertical for the feed, 16:9 for the web, 21:9 for something cinematic. Start short and low-resolution while you are still testing; the credit counter tells you what each choice costs.
Review the result and edit a single variable — the lens, the light, the verb. Two or three passes usually land it, and you learn which word did the work.
It is strongest exactly where there is nothing to shoot and no time to build.
City skylines, weather, interiors, abstract texture — the connective footage that eats a stock-library budget. Generate exactly the shot the edit needs instead of settling for the closest match.
Before anything is designed or shot, generate five interpretations of a brief and put them in front of people. Arguing over moving pictures is far more productive than arguing over adjectives.
A daily posting cadence is impossible with a shoot schedule and practical with a prompt. Vertical ratio and synchronized audio help finish the clip at download; commercial use depends on your paid-plan status and the Terms of Service.
It generates a video clip from a written description alone, with no uploaded footage or images. On FrameAI, Seedance 2.0 produces 4-to-15-second clips at up to 4K, with camera movement, lighting and synchronised audio inferred from your prompt.
One or two precise sentences beat a paragraph. Cover subject, action, camera and light; drop words like "cinematic" and "masterpiece", which add no information. The field accepts up to 2000 characters, but you will rarely need a fraction of that.
Yes. Seedance 2.0 understands Chinese and English natively — neither is translated into the other first, so idiom and camera terminology survive intact in both.
Usually the prompt asked for too much at once. Reduce it to a single action with a single camera behaviour, then add detail back one element at a time. Vague qualifiers also hurt — replace them with concrete nouns, lens choices and lighting terms.
Yes, directly in the prompt — dolly, pan, orbit, crane, handheld, locked-off are all understood. For frame-accurate control over a specific move, the Motion Control tool exposes it explicitly instead.
As many as your credits allow — there is no daily cap. Because cost scales with resolution and duration, drafting at 480p on Seedance 2.0 Mini lets you iterate many times for the price of one finished 4K take.
Each tool is the same Seedance 2.0 model pointed at a different job. Move between them freely — your credits, history and exports are shared.
Write one line, pick a ratio, press generate. No editing suite, no render farm, no post-production pass — a finished clip with sound, ready to download.