Thirty seconds in one take

Thirty Seconds Seedance 2.5 Generates in One Take

Twice the ceiling of the 2.0 series, with up to 30 reference inputs and audio generated alongside the picture rather than added after.

Seedance 2.5, and the Five Models You'd Compare It To

We ran the same prompt through all six. Every one of them is in the box above.

“A chef in a busy kitchen looks straight at camera and says, "the secret is butter."”
Only some of these models generate sound with the picture.
  • Seedance 2.5 result for the speech promptOn this page

    Seedance 2.5

    Native audio, generated with the picture.

    Run this prompt

  • MiniMax H3 result for the speech prompt

    MiniMax H3

    Mouth shapes against the audio track.

  • Kling 3 Omni result for the speech prompt

    Kling 3 Omni

    Skin and hair detail as the camera settles.

  • Veo 3.1 result for the speech prompt

    Veo 3.1

    Whether the gaze holds the lens.

  • Happy Horse 1.0 result for the speech prompt

    Happy Horse 1.0

    How the face holds when the head turns.

  • Wan 2.5 result for the speech prompt

    Wan 2.5

    How skin and cloth read up close.

Seedance 2.5 video generation on Diffus
Behance logoGoogle logoCanva logoSquarespace logoAmazon logo
Measured against our own lineup

Why Seedance 2.5 Stands Out

Every figure below is the model's actual envelope in our generator, compared with the seventeen models beside it.

Thirty Seconds in One Generation

Twice the ceiling of the 2.0 series — anything from 4 to 30 seconds, a second at a time, instead of stitching short clips together.

Generate with Seedance

Thirty References, Ten Audio, Ten Video

Multi-reference mode takes up to 30 images, 10 audio tracks and 10 videos in one shot — where Seedance 2.0 caps at 9, 3 and 3. Enough to hold a cast and a look across a scene.

Generate with Seedance

Cinematic 21:9, and Audio Included

Six aspect ratios from 21:9 down to 9:16, with the audio track generated alongside the picture rather than added after.

Generate with Seedance

Core capabilities

What You Can Build With It

The controls the generator exposes for this model.

Thirty seconds

A Whole Scene, Not a Fragment

Thirty seconds is long enough for a shot to breathe — an establishing move, a beat, a resolution — without cutting between separate generations that never quite match.

Choosing a video model on Diffus
Thirty references

Bring Your Cast and Your Look

Feed in up to 30 reference images plus audio and video references, and keep characters, wardrobe and grade consistent across the whole clip.

Directing motion in a generated clip
Native audio

Sound That Ships With the Picture

Audio is generated with the video rather than laid over it afterwards, so the clip is usable the moment it finishes.

A generated clip with its audio track