All models

kling-v3-omni ★

KuaishouVideo
Get your API key
kling-v3-omni

Controlled Short Video Creation Driven by Multi-Image and Video References

kling-v3-omni is Kuaishou Kling's video model for reference-based creation, also known as Kling 3.0 Omni. It can generate short videos from text or a first-frame image, and can also create and edit using multiple images, reference videos, or existing footage. Its value lies in incorporating character, scene, and style assets into a single task, making it suitable for short-form video production with existing visual assets that requires repeated adjustments to shot content.

KuaishouModel Brand
VideoModel Type
VideoTask Capability

Specifications and API Features

Clarify capacity, inputs and outputs, and invocation methods before choosing a model.

Creation Methods
Text-to-video, image-to-video, multi-image reference, video reference, and existing video editing
Generation Duration
This platform entry: integer durations from 3–15 seconds
Aspect Ratios
16:9、9:16、1:1
Generation Modes
Standard generation: std / pro / 4k; Omni reference requests do not use 4k
Number of References
Up to 7 images without video; up to 4 images and 1 video with video
Audio Control
Standard generation can enable audio; video references can preserve original audio and do not generate new audio simultaneously
Task Output
Video link, video ID, task ID, duration, and status; supports asynchronous processing

The above values and operating ranges apply to this platform's invocation entry. Standard generation and Omni reference editing use different parameter combinations.

Core Capabilities

Learn what kling-v3-omni can bring to your work.

Turn visual assets into creative references

Multi-image references mean that characters, scenes, and styles do not all need to be restated in text. You can place subject photos, environment images, and style images into the same task, then clearly reference them in the prompt and explain their respective purposes. This is especially suited to creation with existing brand assets, allowing shots to develop around clear visual references rather than starting from scratch every time.

Distinguish between editing footage and reference footage

Video references offer two different uses: base treats the video as a foundation to be edited, allowing changes to elements, composition, color, weather, or overall style; feature extracts characteristics such as style and camera movement to guide a new video. Before choosing, clarify whether you want to change the original footage or draw on its characteristics, which helps organize assets and prompts.

Organize motion from first and last frames

Image-to-video uses the first frame as a starting point and can also include a last frame to constrain the ending image, making it suitable for short videos with existing storyboards or key visuals. Standard generation lets you choose the aspect ratio, duration, and whether to generate accompanying audio; when you need to combine reference assets, switch to the Omni workflow and use assets and text together to describe subject actions and scene changes.

Use Cases

Start with specific tasks to find where the model can be effective.

Product and brand short videos

Input product images, brand environment images, and shot descriptions to generate short videos for social content or advertising concepts. Prompts should describe the product's location, display actions, and lighting atmosphere, while clearly defining the role of each reference image. After delivery, focus on checking the product's appearance and visual details, then adjust different creative versions around the same set of assets.

Creative revisions of existing footage

Use an existing video as base input and describe the style, background elements, or weather you want to change, such as turning live-action footage into an anime visual. This is suitable for creating creative variations based on original footage rather than rebuilding an entire storyboard. Audio can be kept or removed, with music, subtitles, and final editing added after the visual edits are complete.

Reference-driven storyboard exploration

Set a reference video as feature and pair it with character or scene images to explore new videos with specified visual styles or shot characteristics. This is suitable for director previsualization, storyboards, and series content concepts. For each task, first focus on one shot objective; after delivering the short video, compare composition, action, and style before deciding the creative direction of subsequent shots.

How to choose this model

Choose based on task complexity, input materials, and expected results.

How to choose between it and kling-v3

If the task mainly relies on text, first and last frames, and explicit camera parameter controls, kling-v3 is better suited to this workflow; if you need multi-image combinations, reference videos, or direct editing of existing shots, prioritize kling-v3-omni. Both can generate short videos with audio, but Omni does not support camera_control, so reference camera movement and camera parameter controls should not be conflated.

How to choose between it and kling-o1

kling-o1 also supports omni references and is not limited to processing a single image. Specific reasons to choose kling-v3-omni include needing more flexible generation durations, audio for standard generation, or processing longer reference videos. Existing O1 workflows can retain similar material organization approaches, but when switching models, you should reset the duration and audio combination rather than simply replacing the name.

Get started

From a small-scale task to formal integration.

01

Prepare the task and materials

Define the goal, required inputs, and output requirements, using real business examples as a starting point.

02

Try it in the API debugging area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate according to the API documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage limits

Before formal use, understand the output quality and capability scope.

  • Omni reference requests cannot use 4k, negative_prompt, cfg_scale, or camera_control. When visual adjustments are needed, express them through reference materials and positive prompts; standard generation can optionally use 4k, but this does not mean reference video editing can use the same mode.
  • When a reference video is included, generate_audio must be false; retaining the original sound is controlled by keep_original_sound. Base editing videos cannot also specify first and last frames. The last frame for image-to-video must also be used together with the first frame; you cannot upload only the ending image.
  • Reference videos are limited to MP4/MOV, under 200MB, 24–60fps, and 3–15.5 seconds, and must also meet dimension and total pixel constraints. Materials must be referenced in the prompt; mixing first and last frames with reference images affects material order, so organize them consistently before submitting.

Frequently Asked Questions

Answers to common questions about using kling-v3-omni.

What is the relationship between Kling O3 and kling-v3-omni?

Kling O3 and Kling 3.0 Omni are names for this model; this platform uses kling-v3-omni when calling it. It corresponds to different models than kling-v3 and kling-o1, respectively. When creating a task, explicitly specify the target model to avoid relying on the default selection.

How should I choose between multi-image references and image-to-video?

When you already have a defined starting frame, use image2video and provide the first frame; add an end frame if necessary. When you need to combine character, environment, and style assets, use text2video with image_list. Reference images must not only be uploaded, but also explicitly referenced in the prompt, explaining their purpose in the shot.

Can I edit videos and also use videos as references to generate new clips?

Yes, but the assets have different roles in the two cases. refer_type=base means editing this base video; feature means drawing on its style, camera movement, and other characteristics to guide creation. The former is suitable for modifying existing content, while the latter is suitable for reference-driven new shots; the two tasks should use different prompt objectives.

Can 4K and audio be used together in all tasks?

No, the same combination does not apply to all tasks. Standard generation supports 4k and audio options, while Omni reference requests do not use 4k; after adding a reference video, new audio generation must be disabled. If you want to preserve the source audio, you can set keep_original_sound instead of enabling generate_audio.

How do I retrieve the generated video after submitting a task?

Submit the model, operation, and creation input to POST /kling/videos. Set async=true to first obtain a task_id and then query the task; you can also use callback_url to receive the result. After completion, retrieve the video through video_url, and save the task ID, status, and asset configuration for convenient tracking and iteration.

Model information · Updated: 2026-10-01. For request parameters and billing rules, see the API and pricing sections.

Use kling-v3-omni for your next task

Start with a clear goal and evaluate whether it suits your work based on real results.