Kling 2nd-Generation Video Model for Short-Shot Creation
kling-v2-master is the V2 Master video model in Kuaishou's Kling series, designed for short-shot creation starting from text concepts or a first-frame image, with a balance between visual quality and user experience. It is suitable for product showcases, concept storyboards, and animated photo assets, and can also handle animation tasks in talking-photo workflows. When choosing, focus on single-shot expression rather than multi-asset editing or native audio-video generation.
Clarify capacity, inputs and outputs, and invocation methods before selecting.
Creation methods
Text-to-video, first-frame image-to-video
Video duration
Video generation on this platform: 5 or 10 seconds
Aspect ratio options
16:9, 9:16, 1:1
Model control scope
Single mode; end frames and camera_control are not supported
Audio workflow
Video generation does not support synchronized audio; talking photos use external audio
Photo voiceover input
Photo URL and audio URL; supported audio formats: mp3/wav/m4a/aac, ≤5MB
Results and tasks
Returns a video link, task ID, and status; supports asynchronous processing and callbacks
The above duration, aspect ratio, and asset requirements apply to this platform's invocation scope. Talking photos are a combined creation workflow and are not equivalent to the model's native audio.
Core Capabilities
Learn what kling-v2-master can bring to your work.
From Text Concepts to Short Shots
Text-to-video is suitable for turning scene descriptions into watchable shot drafts. When creating, you can organize prompts around the subject, action, environment, and lighting, focusing a piece of content on one clear event. The balanced positioning of V2 Master is suitable for comparing different creative expressions, then using selected results for storyboard discussions or subsequent editing.
Use the First Frame to Establish a Visual Starting Point
Image-to-video uses an existing image as the opening foundation, then describes the desired action through prompts. Compared with starting entirely from text, this approach lets product images, portraits, or scene designs participate directly in creation. It emphasizes developing the image from a specified starting point and does not provide end-frame locking, making it more suitable for short shots where the ending is allowed to develop freely.
Photo Animation and Voiceover Integration
The talking photo workflow accepts a portrait photo and existing audio, first animates the photo, and then processes lip synchronization according to the audio. After specifying V2 Master, it handles the animation generation portion. This workflow is suitable for character introductions or short spoken presentations with existing voiceovers, without mistakenly writing dialogue as instructions for the video model to generate sound automatically.
Use Cases
Start with specific tasks to find where the model can be effective.
Create Dynamic Assets from Static Product Images
Input a product display image and describe a slight rotation, background atmosphere, or presentation action to generate dynamic clips for short-video editing. You can choose a landscape, portrait, or square aspect ratio based on the placement, then add selling-point captions and music during editing. It is recommended to highlight one display focus at a time rather than arranging an entire advertising storyline within the same shot.
Concept Storyboards and Shot Previsualization
Organize an individual scene from a script into descriptions of the subject, action, and environment, and use text-to-video to create shot drafts; when art direction already exists, you can also start from a scene image. The deliverables are short clips that facilitate team discussion, suitable for comparing mood, composition, and action expression before deciding which content enters formal production or manual refinement.
Short Character Videos with Existing Voiceovers
Prepare a clear front-facing photo of one person and a voiceover, then use the talking photo entry to create character introductions, character greetings, or brief explanations. The audio should fit within the selected video duration, and prompts should mainly describe expressions and movements. Once complete, obtain the final video link and integrate it into existing content publishing or editing workflows.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
How to Choose Between It and V2.1 Master
V2 Master is suitable for short-shot tasks driven by text or a first frame; V2.1 Master places greater emphasis on quality and consistency. Both use a single mode here, with durations of 5 or 10 seconds, and neither provides an end frame, audio accompaniment, or structured camera movement. If visual consistency is important, compare results using the same materials rather than treating version numbers as a direct guarantee of results.
When to Choose Other Kling Versions
If the task must lock the ending image, consider V2.5 Turbo pro, which supports end frames; if synchronized audio accompaniment is needed, choose V2.6 pro or V3. Multiple images, reference videos, and existing video editing are better suited to O1 or V3 Omni. The value of choosing V2 Master lies in its clear text-to-video and first-frame creation workflow, not in covering these different control requirements.
Get Started
From a small-scale task to formal integration.
01
Prepare the Task and Materials
Define the goal, required inputs, and output requirements, and use real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to view the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Before formal use, understand the output quality and capability scope.
V2 Master does not support end_image_url end-frame constraints or camera_control camera movement parameters. Prompts can express desired actions or camera feel, but this is not the same as structured control and cannot guarantee that the final image stops at a specified composition.
Video generation does not support generate_audio synchronized audio accompaniment, nor does it support 4K mode. Talking photos require audio to be provided separately, and their sound comes from the input material; if dialogue, ambient sound, and visuals need to be generated together, choose a model with the corresponding capabilities.
Do not treat image_list or video_list as multi-material creation capabilities for this model. Reference video editing and Omni multi-image reference belong to other models; V2 Master's image-to-video workflow starts from a first-frame image, so complex material combinations should be produced separately or use another model.
Frequently Asked Questions
Answers to common questions when using kling-v2-master.
Which model name should I enter when making a call?
Explicitly specify model=kling-v2-master in /kling/videos or /kling/talking-photo. This clearly selects V2 Master instead of relying on the default value of the endpoint. kling-v2-1-master is another model and should not be used as an alternative name for the same model.
What core inputs are required for image-to-video?
Use /kling/videos, set action to image2video, provide start_image_url, and use prompt to describe the action and scene changes. Simply choose a duration of 5 or 10 seconds. This model does not support end-frame control, so when creating, treat the first frame as the starting point rather than constraining both the start and end points.
Can a photo directly speak the lines in the prompt?
You cannot use a video prompt as speech generation input. To create a talking photo, provide image_url and audio_url to /kling/talking-photo; prompt is used for movements or expressions during the photo animation stage. It is recommended to prepare the voiceover first, then have the photo lip-sync to that audio.
Can videos of any length be generated?
This model generates videos of 5 or 10 seconds, and talking photos also offer these two duration options. Longer content can be split into shots and then edited together; if you need more flexible single-generation durations, consider V3 rather than passing arbitrary seconds to V2 Master.
How do I integrate generation tasks into an application?
You can use async=true to obtain task_id, then query the task result; you can also set callback_url to receive completion notifications. Applications should associate the task ID with the creation record, then read state and video_url after completion. Successful task submission and completed video generation are different stages and should not be treated as the same status.
Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.