Best Text to Video Prompts for Seedance 2.5 (With Real Outputs)
Seven Seedance 2.5 prompts fired on a fixed seed, with every output embedded and the exact billed cost of each clip.
Most of the bad Seedance 2.5 clips I have generated were not the model's fault. They were prompts that named a subject and then went quiet: no camera, no light, no sound, nothing that told the model what kind of shot it was supposed to be making. Seedance 2.5 fills those gaps itself, and what it fills them with is rarely what you had in mind.
So I fired eight text-to-video prompts through Seedance 2.5 on Segmind, all on a fixed seed, and kept every output. Each one demonstrates a single prompt pattern you can lift and reuse. Two of them are the same scene at the same seed and duration, differing only in how the prompt is written, which is the clearest look at what structure actually buys you. Every clip is embedded below and every cost is the real billed figure from the API response, not an estimate. Total spend for the set: $15.65.
How Seedance 2.5 reads a prompt
Before the patterns, the mechanics that decide what the prompt has to carry. Seedance 2.5 is a synchronous endpoint: you POST a JSON body, the connection stays open while the video renders, and the response body is the MP4 itself. There is no job to poll.
Three defaults matter more than the rest:
aspect_ratiodefaults toadaptive, which lets the model pick the geometry from your prompt. For text-to-video, set it explicitly:16:9,9:16,1:1,4:3,3:4or21:9.generate_audiodefaults totrue, so the model is scoring sound whether or not you wrote anything about it. An audio clause is not an extra, it is the difference between sound you chose and sound you got.durationis any whole number of seconds from 4 to 30 at 24fps, or-1to let the model choose. Duration is also most of what you pay for, so it is a budget decision as much as a creative one.
The prompt does the rest, and it has five slots worth filling: subject, action, camera, light, audio. Everything below is a variation on filling those well.
Pattern 1: write the shot, not the subject
Here is the whole argument for structured prompting in one pair of clips. Same scene, same seed, same duration, same resolution. The only difference is the prompt. First, the version most people type:
Parameters duration: 6s | resolution: 720p | aspect_ratio: 16:9 | generate_audio: true | seed: 42 | billed: $1.431585
Six words of prompt. One continuous pull-back the model chose by itself.
This is not a bad clip, and that is the point worth sitting with. Given almost nothing, Seedance 2.5 invented a professional stainless kitchen, a black apron and a toque, tossed peppers, a blue gas flame, and a slow continuous pull-back from a close-up on the pan to a medium of the chef. Scene detection finds no cuts in it at all: it is one move, competently executed.
What it is not is your shot. The light is flat overhead white. The subject you named spends four of six seconds out of frame, because the model opened on the pan and only arrived at the person at the end. Nothing here was your decision.
Now the same idea with the five slots filled. Nothing clever, just direction that was previously missing:
Parameters duration: 6s | resolution: 720p | aspect_ratio: 16:9 | generate_audio: true | seed: 42 | billed: $1.431585
The same scene with subject, action, camera, light and audio all specified.
Four of the five instructions land exactly. The key light is warm and comes through a window on the left, as asked, and the whole grade shifts with it. The apron is white. The action plays as a complete three-part beat: lift, one toss, pan back down on the flame. The background falls away into shallow focus. Still no cuts.
The fifth instruction is the interesting one. I wrote "locked at chest height", and Seedance 2.5 obeyed it literally: the camera sits at the chef's chest and his head is cropped out of frame for most of the clip. That is a faithful reading of what I asked for and not at all what I wanted. Camera instructions are executed literally, so describe the frame in relation to the subject, not in absolute terms. "Holds him from chest to just above the head" would have fixed it.
The wording pattern generalises to every clip below:
- Subject and action: one specific person doing one specific thing, with a beginning and an end. "Lifts the pan, tosses once, sets it back down" is a complete beat in six seconds. "Cooking" is not.
- Camera: name the move, the framing and the height, then say no cuts if you want one continuous take.
- Light: direction and quality, not a mood word. "Warm late-afternoon light from a window on the left" tells the model where to put the key.
- Audio: name the sources you want, then close the set with "no music".
Pattern 2: multi-shot with explicit Shot blocks
Seedance 2.5 can cut inside a single clip, and labelling the beats Shot 1:, Shot 2:, Shot 3: is how you decide where those cuts land instead of discovering them. This is the pattern for a product spot: three beats, one continuous grade, no editor involved.
Give each shot its own subject, action and camera line, then state the grade once at the end so it carries across all of them. Length matters here. Three beats in six seconds gives each one two seconds, which is not enough to read.
Parameters duration: 15s | resolution: 720p | aspect_ratio: 16:9 | generate_audio: true | seed: 42 | billed: $3.564153
Three labelled shots, one grade, 15 seconds. Cuts landed at 6.9s and 13.0s.
All three beats arrive in the right order: the case opening on concrete, the earbud going in at the window, the wide from behind as the city resolves outside. The cool daylight grade holds across all three, which is the part that would cost you a colourist otherwise.
The useful finding is in the timing. Scene detection puts the cuts at 6.9 seconds and 13.0 seconds, so the model split 15 seconds roughly 7 / 6 / 2. The payoff wide, the shot the whole spot is built toward, got two seconds. If a specific beat needs to breathe, say how long it runs ("Shot 3, five seconds:"), or budget extra duration and expect the last beat to be the one that gets squeezed.
Worth noting what it invented while following orders: I asked for a city outside the window and got a specific, recognisable real skyline. The earbuds, on the other hand, carry no invented brand, which I attribute to naming them plainly as "a matte black wireless earbud case". Close the description of anything brand-shaped or it will fill in a logo for you.
Pattern 3: put dialogue in quotation marks
Audio is generated in the same pass as the video, which is what makes spoken lines work at all: the model is scoring the mouth and the voice together rather than dubbing onto a finished picture. Put the line in quotation marks, attribute it to the speaker, and state that the lips are synced to it.
Parameters duration: 10s | resolution: 720p | aspect_ratio: 16:9 | generate_audio: true | seed: 42 | billed: $2.379393
A spoken line, a static medium shot, no cuts. Audio peaks at -6.6 dB, the hottest of the eight clips.
The picture does its half of the job: a static medium at eye level, no cuts, visible speech articulation through the first half, then the smile and the cup sliding forward, all in one unbroken ten-second take.
Keep the line short. Ten seconds is roughly one sentence of natural speech plus a beat either side, and a line that does not fit gets rushed.
Pattern 4: one camera move per clip
The failure mode with no camera direction is not one bad move, it is several: the model invents its own coverage. If you want a single unbroken take, say so in those words and give the move a direction and a constant speed. This is also where 21:9 earns its keep, because a lateral track across a wide frame is exactly what an ultrawide composition is for.
Parameters duration: 10s | resolution: 720p | aspect_ratio: 21:9 | generate_audio: true | seed: 42 | billed: $2.391010
A single unbroken lateral track at 1470x630. Zero cuts, and the props arrive in the order the prompt lists them.
This is the cleanest execution of the eight. One continuous left-to-right track at constant speed, bench height, zero cuts, landing exactly on the hands with the block plane. Cool light from high windows with real atmosphere in the air.
The lesson hiding in it: the order you list things becomes the order the camera passes them. Hand tools, then shavings, then the half-built chair, then the hands. Write the prompt as the dolly path and the model reads it as one.
One pricing note, since 21:9 is the one ratio that is not free. The file came back 1470x630, which is 0.488% more pixels than 1280x720, and it billed 0.488% more than the same clip at 16:9 would have. Cost tracks output pixels to three decimal places, so ultrawide is a rounding error rather than a premium tier.
Pattern 5: frame for vertical, do not just set 9:16
Setting aspect_ratio: "9:16" changes the canvas. It does not change the composition the model would otherwise have chosen, and a landscape idea cropped to a phone is a waste of the format. Say what vertical means for this shot: how much of the subject is in frame, and where the empty space for on-screen text goes.
Parameters duration: 8s | resolution: 720p | aspect_ratio: 9:16 | generate_audio: true | seed: 42 | billed: $1.905489
A true 720x1280 vertical with headroom left for a text overlay.
The beat plays as written: steps off the kerb, hits the puddle, looks up into the rain. The headroom for an overlay is there. The framing came out a little wider than the "head to knee" I asked for, so the subject sits smaller in frame than intended, which is the same literalism problem as the chef clip in reverse.
One thing to build into any vertical workflow: check a frame, not just the dimensions. On an earlier run I had a 9:16 request come back as a correct 720x1280 portrait file with landscape content rotated sideways inside it, no error and no rotation metadata to warn me. It has happened once in three vertical fires now, this clip being fine, so it is intermittent rather than systemic. Nothing in the response tells you which one you got.
Pattern 6: name every sound, then close the set
The audio clause deserves the same specificity as the camera clause. Name the sources, place them near or far, and finish with "no music". That last part is not stylistic: asking for music by genre can trip a copyright check on the finished audio track and lose you the clip after it has already rendered. Named diegetic sound does not.
Parameters duration: 8s | resolution: 720p | aspect_ratio: 16:9 | generate_audio: true | seed: 42 | billed: $1.905489
Named sources only, no music. Faithful picture, and the quietest track of the eight at -34.9 dB mean.
The picture is faithful down to the details: rain visibly falling, the plastic sheet going over the stacked fruit, the single bare bulb, one static wide with no cuts. It rendered citrus rather than the mangoes I asked for, which is the kind of substitution worth knowing about if the specific product matters to you.
Pattern 7: draft at 480p, then re-fire at 720p
The last pattern is a workflow one, and it comes with a caveat I did not expect. Wording reads perfectly well at 480p, which is much cheaper than 720p for the same seconds. Here is the structured chef prompt again, identical in every respect except resolution:
Parameters duration: 6s | resolution: 480p | aspect_ratio: 16:9 | generate_audio: true | seed: 42 | billed: $0.636754
The same prompt and the same seed at 480p. Note it is a different take, not a preview.
The cost case is strong: $0.64 against $1.43 for the identical six seconds, a saving of 55.5%, and the two prices track the pixel counts almost exactly. Iterating a prompt six times at 480p costs less than four fires at 720p.
The caveat is that a fixed seed does not carry across resolutions. Same prompt, same seed: 42, and the 480p render is a different take: a different kitchen, a different camera path, and this time the chef's head stays in frame for the whole clip. What did transfer was everything the prompt actually specified, the warm window key, the white apron, the toss, the pan returning to the flame. So draft at 480p to test wording, and expect a fresh performance when you re-fire at 720p. It is not a preview.
What the eight clips cost and how long they took
Billed figures straight from the x-cost header on each response, at the published rate of $10.97 per million tokens for 480p and 720p:
| Clip | Config | Billed | Render time |
|---|---|---|---|
| 480p draft | 6s, 480p, 16:9 | $0.636754 | 353s |
| Bare control | 6s, 720p, 16:9 | $1.431585 | 345s |
| Structured | 6s, 720p, 16:9 | $1.431585 | 174s |
| Vertical | 8s, 720p, 9:16 | $1.905489 | 455s |
| Audio design | 8s, 720p, 16:9 | $1.905489 | 529s |
| Dialogue | 10s, 720p, 16:9 | $2.379393 | 408s |
| Camera move | 10s, 720p, 21:9 | $2.391010 | 303s |
| Multi-shot | 15s, 720p, 16:9 | $3.564153 | 360s |
Two things to take from that table. Cost is perfectly predictable: it is a function of seconds and output pixels, nothing else, and generate_audio is free. Render time is not predictable at all. These eight ranged from 174 to 529 seconds with a median near 350, fired two at a time, and a 15-second clip came back faster than a 6-second one. Budget by cost, never by wall clock.
All eight fires succeeded on the first attempt, with no content refusals. Every audio clause in this run ended in "no music", which is the single cheapest insurance you can buy against the output-side copyright check.
Calling it from code
One POST, binary response, write it to disk:
import requests
url = "https://api.segmind.com/v1/seedance-2.5"
data = {
"prompt": "A chef in a white apron lifts a steel pan of sizzling vegetables ...",
"duration": 6,
"resolution": "720p",
"aspect_ratio": "16:9",
"generate_audio": True,
"seed": 42,
}
r = requests.post(url, json=data, headers={"x-api-key": "YOUR_API_KEY"}, timeout=1800)
if r.status_code == 200:
with open("out.mp4", "wb") as f:
f.write(r.content)
else:
print(r.status_code, r.text)
Two things I would build in from the start. Set the timeout high, because the render happens inside this request. And log the x-cost response header on every call: that is the actual amount the generation billed, and it is the cleanest per-call spend record you get. If you are firing a batch, keep the calls sequential rather than threading them out of one process.
Honest assessment
Three things will bite you, and none of them are prompt problems.
Wall clock is unpredictable. Cost is deterministic to the cent; render time is not. The spread in this run was 3x on comparable configs, and on previous runs the same configuration has returned in as little as 80 seconds. Never promise a completion time in a user-facing flow.
There is a moderation pass on the finished video, not just on the prompt. I have had a clip render fully and then come back as a 400 citing possible copyright restrictions, on a prompt that named no title, studio or person. Genre vocabulary was the trigger: words like "detective" and "anamorphic" in a rain-and-neon scene were enough, and dropping them let the identical scene through. The refusal bills nothing, so it costs you the render time and nothing else.
Literal obedience cuts both ways. The chef clip lost its subject's head to a camera height I specified carelessly. The model is not going to second-guess an instruction that produces an awkward frame, so read your camera clause back and ask what it would mean to someone following it exactly.
FAQ
What makes a good Seedance 2.5 prompt?
Fill five slots: subject, action, camera, light and audio. Subject and action are what most people write; the camera, light and audio clauses are what stop the model from making those choices for you. One camera move per clip, named sound sources and an explicit aspect ratio cover most of the gap.
How long should Seedance 2.5 prompts be?
Long enough to fill those five slots, which puts most prompts between 60 and 120 words. Past that you are usually adding adjectives rather than direction. Multi-shot is the exception: each "Shot N:" block needs its own subject, action and camera line.
Does Seedance 2.5 generate audio from the prompt?
Yes. generate_audio defaults to true, the audio is scored in the same pass as the picture, and it costs nothing extra. Naming your sound sources is the only way to control what you get. Expect quiet output: the clips here averaged about -30 dB, so plan on normalising.
How much does a Seedance 2.5 clip cost?
It is token-based, billed on the output video at $10.97 per million tokens for 480p and 720p and $11.99 for 1080p. Measured here: six seconds at 720p billed $1.43, the same prompt at 480p billed $0.64, and 15 seconds at 720p billed $3.56. Resolution and duration are the only levers.
Can I get the same clip twice from the same prompt?
Fix the seed, as every prompt in this post does, and hold the duration and resolution too. The seed does not survive a resolution change: the same prompt and seed at 480p and 720p gave two different takes.
What aspect ratios and durations does Seedance 2.5 support?
Aspect ratios are 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 and adaptive, which is the default. Duration is any whole number of seconds from 4 to 30 at 24fps, plus -1 to let the model decide. Resolution is 480p, 720p or 1080p.
Where to take this next
The pattern behind all eight is the same: decide the shot yourself and write it down, because anything you leave out is a decision the model makes for you. Draft at 480p while you are still editing words, hold the seed, and re-fire at 720p once the wording is right.
Every prompt above runs as-is on Seedance 2.5 on Segmind. Copy one, swap the subject, keep the camera, light and audio clauses, and you will be most of the way to a usable clip on the first fire.