ByteDance

Seedance 2.0 Fast

Seedance 2.0 generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it leads on instruction adherence, motion quality, and visual aesthetics.

Unknown 50K context 10,000 tokens output
Video

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

ByteDance

Input Context Window

The number of tokens supported by the input context window.

50K tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

10,000 tokens tokens

Open Source

Whether the model's code is available for public use.

No

Release Date

When the model was first released.

Unknown

Knowledge Cut-off Date

When the model's knowledge was last updated.

Unknown

API Providers

The providers that offer this model. This is not an exhaustive list.

ByteDance

Modalities

Types of data this model can process.

Video

Pricing for Seedance 2.0 Fast

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

maxResponseSize 10,000 tokens

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

ByteDance

Configuration & Parameters

The configurable options currently documented for this model.

Mode

Toggle Group
Default: text-to-video

Image

Image URL

Start image to guide the video generation.

Last Image

Image URL

Last frame image for video continuation.

Reference Images

Image URL Array

Reference image URLs to guide style, characters, or composition.

Reference Videos

videoUrlArray

Reference video URLs (total length must not exceed 15 seconds).

Reference Audios

audioUrlArray

Reference audio URLs (total length must not exceed 15 seconds).

Aspect Ratio

Select
Default: 16:9
16:9 9:16 4:3 3:4 1:1 21:9

Resolution

Select
Default: 720p
480p 720p 1080p

Duration

Select
Default: 5
5 seconds 10 seconds 15 seconds

Supported Request Parameters

Parameters currently listed by OpenRouter or the local catalog for this model.

Mode Image Last Image Reference Images Reference Videos Reference Audios Aspect Ratio Resolution Duration
Community discussion

What people think about Seedance 2.0 Fast

Seedance 2.0 Fast discussions are most active in r/seedance, r/Seedance_AI, r/AtlasCloudAI. Top Reddit threads cluster around benchmark and model-comparison threads.

The strongest match in this snapshot has 229 upvotes and 108 comments.

r/seedance 137 upvotes 101 comments March 23, 2026
Official Seedance 2.0 is finally here!!

Link: [https://dreamina.capcut.com/](https://dreamina.capcut.com/)

Plan Subcription:

1 month standard $33 = 67,360 creds

Price/Credits per 10s video

Seedance 2.0 = 930 creds = $0.48 per video

Seedance 2.0 Fast = 750 creds = $0.39 per video

Total clips in $33 plan (10s)

Seedance 2.0 = 72 videos

Seedance 2.0 fast = 89 videos

edit:

Looks like its still rolling out some countries dont have it yet!

Open Reddit thread

Been messing with Seedance 2.0 for the past few weeks. The first couple days were rough — burned through a bunch of credits getting garbage outputs because I was treating it like every other text-to-video tool. Turns out it's not. Once it clicked, the results got way better.

Writing this up so you don't have to learn the hard way.

\---

\## The thing nobody tells you upfront

Seedance 2.0 is NOT just a text box where you type "make me a cool video." It's more like a conditioning engine — you feed it images, video clips, audio files, AND text, and each one can control a different part of the output. Character identity, camera movement, art style, soundtrack tempo — all separately controllable.

The difference between a bad generation and a usable one usually isn't your prompt. It's whether you told the model \*\*what each uploaded file is supposed to do.\*\*

\---

\## The system (this is the whole game)

You can upload up to 12 files per generation: 9 images, 3 video clips, 3 audio tracks. But here's the catch — if you just upload them without context, the model guesses what role each file plays. Sometimes your character reference becomes a background. Your style reference becomes a character. It's chaos.

The fix: . You mention them in your prompt and assign roles.

Here's what works:

| What you want | What to write in your prompt |
|---|---|
| Lock the opening shot | \`@Image1 as the first frame\` |
| Keep a character's face consistent | \`@Image2 is the main character\` |
| Copy camera movement from a clip | \`Reference 's camera tracking and dolly movement\` |
| Set the rhythm with music | \`@Audio1 as background music\` |
| Transfer an art style | \`@Image3 is the art style reference\` |

The key insight: a handheld tracking shot of a dog park can direct a sci-fi corridor chase. The model copies the
\*cinematography\*
, not the content.

https://preview.redd.it/7wphbndr5umg1.png?width=2860&format=png&auto=webp&s=5179924ce3f98ba751eaf3b70c662c5a35190983

\---

\## The prompt formula that actually works

Stop writing paragraphs. Seriously. The model doesn't reward verbosity — anything over \~80 words and it starts ignoring details or inventing random stuff.

Structure: \*\*Subject + Action + Scene + Camera + Style\*\*

Here's a side-by-side of what works vs. what doesn't:

|Part|✅ Works|❌ Doesn't|
|:-|:-|:-|
|Subject|"A woman in her 30s, dark hair pulled back, navy linen blazer"|"A beautiful person"|
|Action|"Turns slowly toward the camera and smiles"|"Does something interesting"|
|Scene|"Standing on a rooftop terrace at sunset, city skyline behind her"|"In a nice location"|
|Camera|"Medium close-up, slow dolly-in"|"Cinematic camera"|
|Style|"Soft key light from the left, warm rim light, shallow depth of field, film grain"|"Cinematic look"|

\*\*Pro tip:\*\* "cinematic" by itself = flat gray output. You have to spell out the actual lighting recipe. Think of it like telling a DP what to set up, not just saying "make it look good."

Full example prompt (62 words):

\> "A woman in her 30s, dark hair pulled back, navy linen blazer, turns slowly toward the camera and smiles. Standing on a rooftop terrace at sunset, city skyline behind her. Medium close-up, slow dolly-in. Soft key light from the left, warm rim light, shallow depth of field, film grain."

https://preview.redd.it/eif5pvv86umg1.png?width=2942&format=png&auto=webp&s=faa4a7557b09bd4d1b6c1096b3185073b36ef91a

\---

\## Settings — the stuff most people skip

\*\*Duration:\*\* Start at 4–5 seconds. I know the temptation is to go straight to 15 seconds, but longer clips amplify every problem in your prompt. Lock in the look first, then scale up.

\*\*Aspect ratio:\*\* 6 options. 9:16 for Reels/Shorts/TikTok. 16:9 for YouTube. 21:9 if you want that ultra-wide cinematic bar look.

\*\*Fast vs Standard:\*\* There are two variants — Seedance 2.0 and Seedance 2.0 Fast. Fast runs 2x faster at half the credits. Same exact capabilities (same inputs, same lip-sync, same everything). I use Fast for all my drafts and only switch to Standard for the final keeper. Saves a ton of credits.

https://preview.redd.it/hj3vdzbj6umg1.png?width=1392&format=png&auto=webp&s=5e0cb61ad98c6ae14cceee93a7767417eb02cedd

\---

\## 6 mistakes that burned my credits (so yours don't have to burn)

\*\*1. Too many characters in one scene\*\*
Three or more characters = faces drift, bodies warp, someone grows an extra arm. Keep it to two max. If you need a crowd, make them blurry background elements.

\*\*2. Stacking camera movements\*\*
Pan + zoom + tracking in one prompt = jittery mess that looks like a broken gimbal. One movement per shot. A slow dolly-in. A gentle pan. Or just lock it static.

\*\*3. Writing a novel as a prompt\*\*
Over 100 words and the model starts cherry-picking random details while ignoring the ones you care about. If your prompt doesn't fit in a tweet, it's too long.

\*\*4. Uploading files without \*\*
This was my #1 mistake early on. Uploaded a character headshot and a style reference, didn't tag them. The model used my character as a background texture. Always assign roles explicitly.

\*\*5. Expecting readable text\*\*
On-screen text comes out garbled 90% of the time. Either skip it entirely or keep it to one large, centered, high-contrast word. Multi-line paragraphs are a no-go.

\*\*6. Fast hand gestures\*\*
"Rapidly gestures while counting on fingers" → extra fingers, fused hands, nightmare anatomy. Slow everything down. "Gently raises one hand" works. Anything fast doesn't.

\---

\## The workflow I use now

After a lot of trial and error, this is what I've settled on:

1. \*\*Prep assets\*\* — Gather a character headshot (front-facing, well-lit), a style reference, maybe a short video clip for camera movement. Trim video refs to the exact 2–3 seconds I need.
2. \*\*Write a structured prompt\*\* — Subject + Action + Scene + Camera + Style. Under 80 words. u/tag every uploaded file.
3. \*\*Draft with Fast\*\* — Run 2–3 quick generations on Seedance 2.0 Fast. Change one variable per run. Lock in the look.
4. \*\*Final render\*\* — Switch to standard Seedance 2.0 for the keeper. Set target duration and aspect ratio. Done.

The whole process takes maybe 5–10 minutes once you know what you're doing.

https://preview.redd.it/u1wy708x6umg1.png?width=2506&format=png&auto=webp&s=f25f95c1d2c0111ec8fb809c05af76582603fca5

\---

\## Some smaller tips that helped me

\- \*\*Iterate one variable at a time.\*\* If you changed the prompt AND swapped a reference AND adjusted duration, you won't know which one caused the improvement (or the regression).

\- \*\*Front-facing headshots for character refs.\*\* Side profiles, group shots, and stylized illustrations give the model way less to work with.

\- \*\*One style, one finish.\*\* "Wes Anderson color palette with film grain" → great. "Wes Anderson meets cyberpunk noir with anime influences" → the model has no idea what you want.

\- \*\*Trim your video references.\*\* Don't upload 15 seconds when you only need 3 seconds of camera movement. Cleaner input = cleaner output.

\---

\## TL;DR

\- Seedance 2.0 is a reference-driven conditioning engine, not just text-to-video
\- Use to assign explicit roles to every uploaded file
\- Prompt formula: Subject + Action + Scene + Camera + Style (under 80 words)
\- Use Seedance 2.0 Fast for drafts (half cost, 2x speed), Standard for final renders
\- Max 2 characters per scene, one camera move per shot, no fast hand gestures
\- Start with 4–5 second clips, then scale duration once the look is locked

Hope this saves someone a few wasted credits. Happy to answer questions if you've been hitting specific issues.

Open Reddit thread
r/AtlasCloudAI 82 upvotes 17 comments April 20, 2026
Seedance 2.0 Fast vs Pro?

Made on [AtlasCloud.ai](http://AtlasCloud.ai)

**Visual quality**

Pro seems to be rather creative under my prompts, and has better light/atmosphere texture. Fast gets you there quicker and has great generation too, for consistency i'd say there's no big difference

**When to use which**

Fast is the obvious choice for iteration. Previs, storyboards, testing prompt ideas, anything where you're trying to figure out if a concept works before committing. Pro is for final and high-quality results.

**Cost**

Depends on providers, nobody seems to have a clean fast vs pro price breakdown that covers every provider. On atlascloud, fast mode is 0.026 cheaper per second than the pro, and the quality is very similar, so I usually go with the fast

>prompt:
Epic wide-angle shot of a vast ancient battlefield at golden hour, thousands of warriors clashing with swords and shields under a hazy amber sky thick with smoke and ash. A lone archer in weathered bronze armor, face streaked with dirt, draws a longbow with deliberate tension. The arrow releases with a sharp twang.
Camera immediately snaps behind the arrow, tracking it in extreme slow motion as it cuts through drifting smoke and falling embers. Shallow depth of field keeps the arrow razor-sharp while the chaotic battlefield blurs behind. The camera pushes closer, tighter, until the wooden shaft fills the frame — revealing intricate carved runes and weathered grain.
Seamless transition to macro scale: the arrow's surface becomes a landscape. A microscopic civilization of tiny warriors the size of splinters wages war across the fletching. Miniature catapults hurl fragments of dust. Warriors scale the carved runes like canyon walls. Torches flicker. Banners wave. All rendered with

Open Reddit thread

Look at this comparison between seedance 2.0 and google veo 3.1 quality:

VEO 3.1:

[VEO 3.1](https://reddit.com/link/1ry153r/video/tu2bzwpza0qg1/player)

Seedance 2.0 Fast:

[Seedance 2.0 Fast](https://reddit.com/link/1ry153r/video/2mjtish3b0qg1/player)

Prompt: [https://pastebin.com/iRX6yHN6](https://pastebin.com/iRX6yHN6)

I am personally completely blown away by how close Seedance 2.0 is to replicating action movies perfectly. Unfortunately the only website I have found to be reliably working with seedance 2.0 at a reasonable price is [yapper.so](https://yapper.so/?via=cliplord) (yes, this is an affiliate link).

I have personally been in touch with the founders of this website, [https://x.com/ehalm\_](https://x.com/ehalm_) and [https://x.com/SeanGrindal](https://x.com/SeanGrindal) and while they are slow to respond at times, their website has been operational since may 2025, and they really do offer the actual seedance 2.0 model.

Open Reddit thread
View more discussions →

More models from ByteDance

Continue browsing adjacent models from the same provider.

← All AI Models