Kling's First-Generation Video Model for Short Shot Creation
kling-v1 is Kuaishou Kling's first-generation video generation model, suitable for turning text concepts or static images into short shots and further extending generated videos. Its practical features include support for five-second first and last frames and camera movement control, making it easy to arrange the start and end of scenes and camera direction. For creations that do not require native audio, 4K, or multi-material editing, this model can be used to build a streamlined video workflow.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation Method
Text-to-video, image-to-video, existing video extension
Generation Duration
5 or 10 seconds, 5 seconds by default
Modes and Aspect Ratios
std / pro; 16:9, 9:16, 1:1
First and Last Frame Control
5-second image-to-video can use a first frame and a last frame; the last frame must be used with a first frame
Camera Movement Control
5-second tasks only; supports camera_control
Prompt Control
cfg_scale: 0–1; negative_prompt: up to 200 characters; extension tasks support neither
Delivery and Photo Lip Sync
Returns a video link and task information; the photo lip-sync feature takes an image and existing audio as input, supporting 5 or 10 seconds
The above are the calling specifications for kling-v1 on this platform; photo lip-sync is a combined creation feature and does not mean the model natively generates audio.
Core Capabilities
Start a Shot from Text or an Image
When no existing visual is available, use text to describe the subject, action, and environment for text-to-video generation; when product or scene images are available, use an image as the first frame for image-to-video generation. The two methods are suited to concept exploration and animating an established visual, respectively, and output video links for continued editing.
Arrange Short Shots with First and Last Frames
Five-second image-to-video generation can specify both the first and last frames, giving the shot a clear starting point and endpoint; five-second tasks can also enable camera movement control. In practice, first determine the composition, then describe the intermediate action, treating shot direction as a separate control rather than relying only on a general prompt.
Continue Expanding After Generation
After obtaining a satisfactory clip, save the returned video_id, use extend, and add the next prompt to continue generation. Extensions are based on an existing generated video and are not equivalent to arbitrary video editing. Applications can use task_id to track tasks and receive video results through asynchronous queries or callbacks.
Use Cases
Turn Product Still Images into Showcase Clips
Provide a product image as the first frame, describe the background atmosphere and desired actions, and create a short showcase shot. When controlling the beginning and ending is needed, use the five-second first-and-last-frame option; when camera movement is needed, also choose a five-second task. The delivered video can be used as advertising edit material, with captions and music added in later production.
Turn Storyboard Concepts into Dynamic Drafts
Write the subjects, scenes, and actions in a storyboard as prompts, then generate five- or ten-second clips to observe the concept in motion. Landscape, portrait, and square aspect ratios can meet different layout requirements. Once a satisfactory shot is available, try extending it and organize the material clip by clip instead of requesting a complete long video at once.
Sync a Portrait Photo to Spoken Audio
Provide a clear front-facing photo of one person and pre-recorded audio, then use the talking photo feature to create a short talking-head video. When calling it, explicitly select kling-v1; prompts are used to describe actions or expressions during the animation stage. The result includes the final lip-sync video, and an intermediate photo animation video can also be obtained.
How to choose this model
Retain the value of choosing v1 when five-second camera movement is needed
If the task focuses on animating still images and requires five-second first-and-last frames or camera movement control, kling-v1 has a clear feature combination. Although kling-v1-6 is a later version, it does not support camera movement and cannot be directly substituted simply because its version is newer. Selection should center on camera control requirements rather than treating the version number as the sole basis for judgment.
Choose other models for audio, 4K, and asset editing
If you need to generate audio while generating video, consider the pro mode of kling-v2-6 or kling-v3; if you need 4K and more flexible durations, consider kling-v3. For multi-image reference or editing existing videos, choose kling-o1 or kling-v3-omni. v1 is better suited to basic creation using text, first frames, and short camera control.
Getting started
Choose a task from text or a first frame
For text shots, select action=text2video; for image animation, select action=image2video and provide start_image_url. Clearly describe the subject's action and the scene, and start with one coherent shot.
Set it to five or ten seconds
Specify model=kling-v1 to /kling/videos, and choose std/pro and 5 or 10 seconds. Five-second image-to-video can use end_image_url, and five-second tasks can use camera_control according to the documentation.
Check silent visuals and continuity
Use async or callback_url to track tasks, and save the task_id and video ID; check the subject, action, and first-to-last-frame continuity. This model does not generate audio, so add sound in post-production; use the corresponding video ID when extending.
Trial suggestion: a five-second first-and-last-frame test
Input and goal
Start from a first frame of a person sitting on a chair and end with a last frame of the person standing up, maintaining the indoor composition, with slow, natural movement and a fixed camera.
Acceptance and next steps
Use 5-second image2video and check the starting and ending frames and body movement; do not reuse five-second-specific last-frame or camera movement controls for 10-second tasks.
Usage Boundaries
Ten-second generation cannot reuse all the controls available for five seconds: kling-v1's first and last frames and camera movement are only applicable to five-second tasks. The last frame can only be used for image-to-video and must be provided together with the first frame; if the shot must end on a specified final image, design it according to the five-second approach first.
kling-v1 does not support native audio or 4K mode, nor does it support Omni multi-image reference or reference video editing. Talking photos require existing audio, using photo animation and lip synchronization, and cannot replace music, dialogue, or ambient sound generation.
Extension tasks do not support negative_prompt or cfg_scale, and cannot reuse all control parameters from the initial generation unchanged. For photo lip-sync tasks, a clear front-facing photo of a single person is recommended; the audio duration should not exceed the selected video duration to avoid mismatches between source material and delivery length.
Frequently Asked Questions
How should I choose between kling-v1's five-second and ten-second tasks?
Choose five seconds when first/last frame or camera movement control is needed; this is a clear functional boundary of this model. Ten seconds is suitable for generation tasks that do not depend on these two controls. The duration should be determined by the shot content; if you need to continue an existing clip, use video extension rather than assuming arbitrary lengths are supported.
Can I provide only the last frame for image-to-video?
No. Image-to-video requires start_image_url; the last frame, end_image_url, is only an optional ending reference and must be used together with the first frame. When using first and last frames with kling-v1, you must also choose a five-second task; when only using the first frame, describe the desired action around that image.
How can I make generated content more closely match the prompt?
For initial generation, cfg_scale can adjust prompt relevance, ranging from 0 to 1; higher values adhere more closely to the prompt. negative_prompt can specify content you do not want to appear, up to 200 characters. Neither applies to extend; when continuing, focus on describing the action in the next segment.
How can I make a photo speak?
Use /kling/talking-photo, submit image_url and audio_url, and explicitly set model to kling-v1. Supported audio formats are mp3, wav, m4a, and aac, with a maximum size of 5MB; videos can be five or ten seconds. It synchronizes lip movements using existing audio and does not create dialogue on its own.
How can generation results be passed to the application for further processing?
Video tasks return information such as video_url, video_id, task_id, and status: the video link is used to retrieve the finished video, the video ID is used for subsequent extension, and the task ID is used to track processing progress. Setting async=true lets you obtain the task ID first and then query the result; you can also configure a callback address to receive completion information.