Seedance 2.5 Media to Video: Combine Image, Video, and Audio References

Sep 22, 2026

Short answer: Give the image authority over the person's appearance, the video authority over movement and camera direction, and the audio authority over pacing or mood. Separate roles reduce conflicts but do not guarantee that no detail transfers between references.

Open the correct workflow first

Open Media to Video, select the model shown by the interface, choose the available reference mode, upload only the assets you need, review settings and estimated credits, and then generate.

At review, Seedance 2.5 Lite had Standard reference enabled. Video edit and Video extend were visible but disabled for that selected model. Visible labels do not establish feature availability; check the selected model and enabled controls before generating. The interface showed up to 9 images, 3 videos, and 3 audio files, with displayed file hints of 30 MB for images, 50 MB for videos, and 15 MB for audio. These limits can change.

Assign one job to each reference

Asset Primary job Do not make it control
Image face, hair, clothing, color, and starting appearance every camera move and cut
Video body movement, gesture, camera direction, or action timing the final subject's identity or background
Audio rhythm, mood, or one clear timing cue exact preservation of the original track

Use a clear person image, a short motion clip with one readable action, and an audio excerpt with one clear accent and a sustained ending. After upload, use the labels shown by the interface, such as @Image1, @Video1, and @Audio1.

Illustrative prompt pattern

This is an illustrative prompt, not a tested result:

Use @Image1 for the person's appearance, @Video1 for the action and camera
movement, and @Audio1 only to guide pacing and mood. Do not copy unrelated
subjects, clothing, or background details from @Video1.

Create a 10-second 16:9 video of the person from @Image1 walking through a
bright studio toward a window. Keep the face, hair, outfit colors, and body
proportions consistent with @Image1. 0–3s: stable medium shot. 3–7s: two
natural steps and a turn toward the window, using the movement language from
@Video1. 7–10s: ease to a stop and hold the final pose as the audio resolves.
Keep one continuous shot with no cuts.

Preserve any markings already visible in @Image1, but add no new text or logos.
Use @Audio1 to guide pacing and mood; exact reproduction of the source audio
is not guaranteed. Keep the same person and outfit throughout. No extra people,
new props, sudden zoom, orbit, or cuts. Preserve natural anatomy and the
intended framing.

Check conflicts before and after generation

Check identity transfer from the motion clip, camera conflicts, multiple audio peaks, and whether the ending is clearly specified. Review appearance and framing first, then motion and camera, then the visible action against the timing cue. If the output is wrong, change one reference or one instruction at a time.

Audio is a timing reference, not an automatic final soundtrack

An audio reference does not mean the original music or narration will be preserved exactly. To retain the source track, import the generated video and source audio into an external editor that displays an audio waveform, align the intended cue with the visible action, and finish the mix there.

FAQ

Does a motion reference copy its subject into the result?

Defining the reference's role can reduce unwanted subject transfer, but it does not prevent it reliably. Check the output for appearance or background details copied from the video reference.

Can audio alone create the visual story?

An audio reference alone does not reliably specify the visual story. Add a visual prompt defining the subject, setting, action, camera, and ending.

Should I upload the maximum number of references?

No. Use the smallest set in which every asset has a distinct job.

Next step

Test one simple reference combination in Seedance 2.5, then change one reference or instruction at a time.

Seedance Team

Seedance Team