Create short videos with sound using text and first and last frames
Seedance 1.5 Pro is ByteDance's video generation model, and doubao-seedance-1-5-pro-251215 is its date-fixed version. It supports text-to-video, image-to-video, and sound generation, making it suitable for turning shot descriptions or static images into short films. When you need to start from a specified image, control the final composition, or add sound to a scene, this version provides a clear creative workflow.
Clarify capacity, inputs and outputs, and invocation methods before choosing a model.
Creation modes
Text-to-video, image-to-video, first and last frame control
Video duration
Platform invocation: 4–12 seconds; duration=-1 for automatic duration
Output resolution
Platform options: 480p, 720p, 1080p
Aspect ratio
16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive
Sound generation
Enable with generate_audio=true; disabled by default
Generation controls
Random seed, fixed camera, watermark options, last-frame return
The values above represent the invocation range for this version on this platform. Creation modes and controls apply to this model and are not equivalent to the capabilities of other Seedance versions.
Core Capabilities
Learn what doubao-seedance-1-5-pro-251215 can bring to your work.
Develop visuals and sound together
1.5 Pro can enable sound while generating video, so you do not need to split every creation into separate visual and audio tasks. Prompts can describe character actions, the environment, and the sounds you want to hear at the same time, such as wind blowing through hair and the sound of wind. Turn off sound generation when you need silent footage, then arrange music or editing later.
Use first and last frames to define shot boundaries
Image-to-video can start from a first-frame image, or include a last-frame image to specify the ending shot, then use text to describe the actions and camera movement in between. It is suited to shot designs with an established composition; unlike providing only character references, first and last frames emphasize the beginning and end of the image rather than identity references across scenes.
Adjust generation around the delivery aspect ratio
Landscape, portrait, square, and ultra-wide formats can serve presentations, mobile short videos, and concept shots respectively. When calling it, you can explicitly set the aspect ratio and duration, and use fixed shots or random seeds depending on the task. When you need to continue planning the next segment of footage, you can request the last frame to be returned as an image for preparing subsequent shots.
Use Cases
Start with specific tasks to find where the model can be effective.
Turn product images into showcase shots
Use a product image as the first frame, add the subject's actions, background changes, and camera direction, and generate short shots for advertising edits. For example, have steam rise from a cup while the camera slowly moves closer, rather than simply asking for it to “look premium.” The deliverable is downloadable video footage, making it easy to add brand text and editing rhythm afterward.
Creative test shots for scenes with sound
Enter a clear short scene, specifying the subject, action, and sound, such as rain falling on a street and a person walking by under an umbrella, then enable generate_audio. This is suitable for first validating creative directions that combine sound and visuals before deciding on subsequent production; after delivery, listen to the audio track rather than checking only individual frames.
Dynamic previsualization between storyboards
When you already have starting and ending storyboards, set the two images as the first and last frames respectively, describe the movement, turn, or camera change between them, and generate a transition preview. It can help directors, animation teams, or design teams discuss whether a shot works, and the delivered video can also be used for presentations rather than replacing full-length production.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Choose 1.5 Pro When You Need Short Videos with Sound
If the task is a text- or first/last-frame-driven short video and you want to generate sound directly, 1.5 Pro is the clear choice. Compared with the 1.0 series, which does not support generate_audio, it adds a path for creating content with sound. If you only need silent visuals, there is no need to enable audio just for the version number; consider the relevant 1.0 models separately for text-to-video and image-to-video tasks.
Choose Other Versions for Complex References and Editing
If you need character reference images, reference audio, or reference video, consider the 2.0 series rather than adding these materials as roles in 1.5 Pro. If you need to edit existing videos, extend clips, or generate up to 30 seconds, choose 2.5; if you need 4k output, consider 2.0 Standard. Your choice should follow the task, not just the version number.
Get Started
From a small-scale task to formal integration.
01
Prepare the Task and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to view the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limits
Understand output quality and capability limits before formal use.
1.5 Pro is suitable for 4–12-second short videos and can also use automatic duration, but it should not be planned for 30-second long-video tasks. Longer narratives can be split into multiple shots and then edited together; returning the last frame helps prepare the next shot, but does not mean existing videos can be extended automatically.
For image input, use first_frame or last_frame. Do not treat reference_image, reference_audio, or reference_video as material modes for this model. Enabling sound generation also does not mean uploading audio references; they correspond to different creation methods.
Each text content item supports up to 1000 characters. Focus on describing the subject, action, shot, and sound. Do not use the frames control supported only by the 1.0 series, and do not set the 2.5-specific edit, extend, or retrieval tools for this model; place parameters in top-level fields whenever possible.
Frequently Asked Questions
Answers to common questions about using doubao-seedance-1-5-pro-251215.
How does 1.5 Pro generate videos with sound?
Set generate_audio=true in the request, and describe the scene and the sounds you want in the text. This option is off by default, so simply writing a sound description without enabling the parameter should not be treated as a sound-enabled delivery solution. After generation is complete, check both the visuals and the actual audio track.
How should first-frame and last-frame images be submitted?
Add image_url items to the content array, put the address in the image_url.url object, and set the role to first_frame or last_frame respectively. First-and-last-frame tasks should also include a text description of the changes in between; image_url cannot be written directly as a string.
Can it generate 4k or 30-second videos?
This version is used with 480p, 720p, 1080p, and 4–12 seconds, and also supports automatic duration. For 4k tasks, consider 2.0 Standard; for tasks up to 30 seconds, consider 2.5. Do not directly use resolution or duration settings from other versions with 1.5 Pro.
Can I use a portrait photo as a character reference?
You can use a portrait photo as the first frame, so the shot begins from that image; however, this is different from reference_image character reference. If the goal is to preserve the person's identity while changing the scene, consider the 2.0 series that supports this reference method. The 1.5 Pro image workflow centers on first and last frames.
How do I obtain the video file after calling it?
Submit model and content to /seedance/videos, with doubao-seedance-1-5-pro-251215 as the model. Asynchronous mode returns a task_id; then query the task, or use callback_url to receive notifications; after the task is complete, download the video through data.video_url.
Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.