The Official Seedance 2.5 Prompt Guide: ByteDance's Six-Part Formula, Explained with Examples
A Seedance 2.5 prompt guide to ByteDance's official six-part formula, @reference binding, bracket syntax for audio, and the parameter locks to know.
Most prompt advice for AI video is folklore. Someone gets a good clip, screenshots the prompt, and a hundred listicles repeat it as a rule. Seedance 2.5 is the first time I have seen ByteDance skip the folklore and publish an actual grammar: a fixed slot order, a tagging syntax for reference files, four bracket types that route audio and on-screen text, and a set of parameters that quietly lock themselves depending on which mode you are in.
That matters because a grammar is testable. You can read it, follow it, and predict what the model will do, which is a different activity from guessing adjectives. Below is the whole system in one place: the six-part formula, the @ reference binding rules, the bracket syntax, how to stage a 30 second clip, and the locked-parameter traps that waste the most credits. Updated 12 August 2026: when this first ran the 2.5 API was not callable anywhere, so the post flagged what was unannounced. That has changed. The model is live on Segmind now, pricing is published, and every rule below is testable against the real endpoint.
What ByteDance actually shipped
Seedance 2.5 raises the single-clip ceiling to 30 seconds and takes up to 50 reference files in one request. The headline capability is not resolution, it is binding: you can hand the model a pile of images, clips and audio, then say in the prompt exactly which file governs which element of which beat. That is what the grammar exists to express.
The access picture is settled as of 12 August 2026: Seedance 2.5 is callable on Segmind at https://api.segmind.com/v1/seedance-2.5, model ID seedance-2.5. Durations run from 4 to 30 seconds in one-second steps, resolution is 480p or 720p, and native audio is a flag rather than a second pass. Billing is token-based on the generated video: $10.97 per million output tokens for text or image input, $6.56 per million when the request includes a reference video, and inputs themselves are free. Per clip at 720p 16:9 that works out to $1.19 for 5 seconds, $2.39 for 10 seconds and $7.17 for 30 seconds. The same 5 seconds at 480p is $0.53.
The prompt grammar is published and stable, and it is the part of the pipeline that survives every version bump. Prompts are portable in a way that pipeline code is not. If you want the operational side, I wrote a companion piece on how to prep your workflow around 2.5. This post is the syntax reference.
The six-part formula
A Seedance 2.5 prompt has six slots, in this order:
Subject + Action or Event + Scene and Environment + Visual Style + Camera Movement or Cut + Audio
Only the first two are required. Everything after "Action" is optional, and this is the part people get wrong in both directions. Some write two words and wonder why the output feels generic. Others write a 300 word paragraph stuffed with every cinematography term they know, which fights itself: the model has to reconcile "handheld documentary" with "locked-off symmetrical composition" and you get mush.
My rule after testing the same structure on 2.0 for months: fill the slots you actually care about and leave the rest empty. Two or three sentences is the sweet spot. Clearer beats longer.
Read that back and you can see why the order helps. Subject and action anchor the shot, scene and style set the look, camera decides how you see it, audio decides what you hear. Each instruction lands in a slot instead of competing for attention.
Binding references with @ tags
This is the real upgrade. Uploaded files get numbered tags you reference directly in the prompt: @Image 1, @Video 1, @Audio 1, and so on. Instead of hoping the model infers what a reference is for, you assign it a job.
The budget per generation:
| Reference type | Hard limit | Range that stays stable |
|---|---|---|
| Images | Up to 30, each under 4K | 1 to 8 distinct subjects |
| Video | Up to 10 clips, 30 seconds combined | 1 to 5 subjects, 5 to 10 seconds each |
| Audio | Up to 10 clips, 30 seconds combined | Keep it to what you need |
Fifty files is the ceiling, not a target. Stability drops as the count climbs, and 30 images of 20 different subjects is a reliable way to get a muddle. Eight or fewer distinct subjects is where results hold together.
The second half of binding is exclusion, and it is the habit most people never build. A reference image carries everything in it: the subject, the background, the lighting, the other people in frame. If you only want the jacket, say so.
@Image 2 defines the courier's face and hair only. Do not use her clothing or the background.
@Video 1 defines the camera rhythm: one slow lateral track, no cuts.
Name every character, product and prop, and tie each one to exactly one file. When two images both plausibly define "the jacket", you have handed the model a coin flip.
The four brackets that control sound and text
Audio and on-screen text are routed by bracket type rather than described in prose. Four brackets, four destinations:
| Bracket | Routes to | Example |
|---|---|---|
( ) | Music and ambient beds | (low cello drone) |
< > | Sound effects | <door latch, rain on glass> |
{ } | Spoken dialogue | {We should go.} |
【 】 | On-screen subtitles | 【We should go.】 |
Two things to note. Dialogue wants its language and delivery declared before the line, not inside it: state "Dialogue language: British English, warm and low", then put only the spoken words inside the braces. And dialogue and subtitles are separate channels. If you put a line in { } and expect it to appear as text on screen, it will not. Repeat it in 【 】 when you want both.
Keep sound effect descriptions short and physical. On 2.0 I found that dense per-object foley requests were far more likely to fail than a couple of concrete sounds, and the same instinct applies here: two or three specific effects beat a paragraph of sound design.
Structuring a 30 second clip in stages
Thirty seconds is long enough that a single flat description falls apart. The recommended shape breaks the clip into stages, each with one job.
[Stage 1] Initial state, one primary event, end state.
[Stage 2] Continue from the previous end state, new event, new end state.
[Stage 3] Closing event and final state.
[Maintain Consistency] What must not change across stages.
The discipline that makes this work: one main change and one clear end state per stage. The moment a stage contains two events, the model has to decide which one to spend its motion budget on, and the transition between stages gets soft. The [Maintain Consistency] block is where you pin down character count, clothing, who is holding which prop, and spacing, which is exactly what drifts across a long generation.
The parameter locks most people trip on
This is the section I wish someone had handed me earlier, because these locks are invisible until you fight them. Depending on the task, aspect ratio and duration stop being yours to set:
| Task | Aspect ratio | Duration |
|---|---|---|
| Video editing | Locked to the input | Locked to input length, give or take about 0.3 seconds |
| First frame, or first and last frame | Locked to the first image | Free to set |
| Video extension | Locked to the input | Free to set |
The practical consequences: you cannot reframe a 16:9 master into 9:16 by running it through an edit, and you cannot use an edit pass to stretch a clip to a target length. If you need vertical, generate vertical, or crop outside the model. And when you extend a clip, expect the appended segment's audio level to not perfectly match the source, so plan on a level pass in the edit.
A checklist before you hit generate
- Does the prompt name the subject and the main action in plain words?
- Does every reference file say both what to use and what to leave out?
- Is every character, product and prop tied to exactly one file?
- Are references introduced scene by scene rather than all forced into frame at once?
- Does each stage of a long clip have one main change and one clear end state?
- Do character count, clothing, prop ownership and spacing stay consistent throughout?
- For an edit, does the prompt name the master video, the exact scope of the change, and what to keep?
Where Segmind fits
Seedance 2.5 is live on Segmind, as promised, and every part of the grammar above has a slot in the request body. The endpoint takes reference_images, reference_videos and reference_audios arrays alongside first_frame_url and last_frame_url, so the @ tagging maps straight onto real fields. duration accepts any integer from 4 to 30, plus -1 for video edits, which is the API expressing the same duration lock the table above describes. generate_audio switches on the track the bracket syntax routes into, and it costs nothing extra.
One caveat before you move a whole pipeline across: 2.5 tops out at 720p. It buys length, reference binding and audio, not pixels. Anything headed for a large screen is still a job for Seedance 2.0, which takes the same reference fields but offers 480p through 4k, caps clips at 15 seconds, and bills output at $7.00 per million tokens against 2.5's $10.97. A higher version number is not a strict upgrade here, so pick per job: length and binding from 2.5, pixels from 2.0. For cheap iteration passes, Seedance 2.0 Fast runs the same reference workflow.
Written in slot order, the same prompt lifts between the two with no edits, which is the useful property here: you pick the model for its resolution and length, not for its prompt dialect.
Honest assessment
A published grammar is a real improvement over vibes, and the exclusion syntax in particular fixes the single most annoying failure mode in reference-driven video. But a formula is not a guarantee. Slot order makes your intent legible, it does not make the model obey, and long generations still drift on the things [Maintain Consistency] is supposed to hold. The 50 file ceiling is also close to a trap: it reads like an invitation and behaves like a cliff. Treat 8 subjects as the real limit and the number in the spec sheet as marketing. And now that the API is open, a sloppy prompt has a price on it: a 30 second 720p miss is $7.17, which makes the checklist above worth the two minutes it takes.
FAQ
What is the Seedance 2.5 prompt formula?
Six slots in order: Subject, Action or Event, Scene and Environment, Visual Style, Camera Movement or Cut, and Audio. Only Subject and Action are required. The remaining four are optional and best left empty unless you specifically want to control them.
Do I need to use all six parts of the formula?
No. Subject and action are the only required slots. Filling every slot with dense detail often makes output worse, because competing style and camera instructions fight each other. Two or three clear sentences outperform a long paragraph.
How many reference files can one Seedance 2.5 prompt take?
Up to 50 total: 30 images, 10 video clips and 10 audio clips. Video and audio are each capped at 30 seconds combined. Stability drops as counts rise, so 1 to 8 distinct subjects is the range that holds together.
What do the brackets mean in a Seedance 2.5 prompt guide example?
Round brackets carry music, angle brackets carry sound effects, curly braces carry spoken dialogue, and the full-width brackets 【】 carry on-screen subtitles. Dialogue and subtitles are separate channels, so repeat a line in both if you want it heard and seen.
Can I change aspect ratio when editing a Seedance 2.5 video?
No. Editing locks aspect ratio to the input and holds duration to roughly the input length, within about 0.3 seconds. First frame generation locks the ratio to your first image but leaves duration free. Extension locks the ratio and leaves duration free.
Is the Seedance 2.5 API available yet?
Yes. As of 12 August 2026 it is live on Segmind as seedance-2.5, with durations from 4 to 30 seconds at 480p or 720p and optional native audio. A 5 second 720p 16:9 clip costs $1.19 and a 30 second one costs $7.17. Seedance 2.0 is still the model to use if you need 1080p or 4k, and it shares the same reference-input workflow.
Wrapping up
The reason to learn this grammar is that it is the cheapest part of the pipeline to get right. Slot order, one job per reference, explicit exclusions, one event per stage, and a clear head about what the locks take away from you. All of it is free to fix in a text editor and expensive to fix in a re-render. Seedance 2.5 is on Segmind now, so you can run the whole grammar against the real model today, and the same prompts still work on Seedance 2.0 the moment you need 1080p or 4k.