Kling's first-generation video model for short-shot creation
kling-v1 is Kuaishou Kling's first-generation video generation model, suitable for turning text concepts or still images into short shots and further extending generated videos. Its practical features include support for five-second first and last frames and camera movement control, making it easy to arrange the beginning and end of scenes and camera direction. For creations that do not require native audio, 4K, or multi-asset editing, this model can be used to build a streamlined video workflow.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Creation methods
Text-to-video, image-to-video, existing video extension
Generation duration
5 or 10 seconds, 5 seconds by default
Modes and aspect ratios
std / pro; 16:9, 9:16, 1:1
First and last frame control
5-second image-to-video can use a first frame and a last frame; the last frame must be used with a first frame
Camera movement control
5-second tasks only; supports camera_control
Prompt control
cfg_scale: 0–1; negative_prompt: up to 200 characters; extension tasks support neither
Delivery and photo lip sync
Returns a video link and task information; the photo lip-sync feature takes an image and existing audio as input and supports 5 or 10 seconds
The above are the invocation specifications for kling-v1 on this platform; photo lip-sync is a combined creation feature and does not mean the model natively generates audio.
Core Capabilities
Learn what kling-v1 can bring to your work.
Start a shot from text or an image
When no existing visual is available, use text to describe the subject, action, and environment for text-to-video; when you already have a product image or scene image, use the image as the first frame for image-to-video. These two methods are suitable for concept exploration and animation based on an established visual, respectively, and output video links for continued editing.
Plan short shots with first and last frames
Five-second image-to-video can specify both the first and last frames, providing a clear start and end for the shot; five-second tasks can also set camera movement controls. In actual creation, you can first determine the composition, then describe the intermediate action, and treat camera direction as a separate control rather than relying only on a vague prompt.
Continue extending after generation
After obtaining a satisfactory clip, save the returned video_id, use extend, and add the next prompt to continue generating. Extensions build on previously generated videos and are not equivalent to arbitrary video editing. Applications can use task_id to track tasks and receive video results through asynchronous queries or callbacks.
Use Cases
Start with specific tasks to find where the model can be useful.
Turn static product images into showcase clips
Input a product first-frame image, describe the background atmosphere and desired actions, and create a short showcase shot. When you need to control the ending, use the five-second first-and-last-frame option; when you need camera movement, also choose a five-second task. The delivered video can be used as advertising editing material, with subtitles and music added in later production.
Turn storyboard concepts into dynamic drafts
Write the subjects, scenes, and actions from a storyboard as prompts, and generate five- or ten-second clips to observe the dynamic presentation of the concept. Landscape, portrait, and square aspect ratios can meet different layout needs. When you already have a satisfactory shot, try extending it and organize material segment by segment rather than requesting a complete long video at once.
Match lip movements to audio for portrait photos
Provide a clear front-facing photo of one person and pre-recorded audio, then use the talking photo feature to create a short talking-head video. When calling it, explicitly select kling-v1; the prompt is used to describe actions or expressions during the animation stage. The result includes the final lip-sync video, and you can also obtain the intermediate photo animation video.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Keep v1 as an option when you need five-second camera movement
If the task focuses on animating still images and requires five-second first and last frames or camera movement control, kling-v1 offers a clear combination of features. Although kling-v1-6 is a later version, it does not support camera movement and cannot be directly substituted simply because it is newer. Choose based on camera control requirements rather than treating version numbers as the sole criterion.
Choose other models for audio, 4K, and asset editing
If you need to generate audio along with video, consider the pro mode of kling-v2-6 or kling-v3; if you need 4K and more flexible durations, consider kling-v3. For multi-image references or editing existing videos, choose kling-o1 or kling-v3-omni. v1 is better suited for basic creation involving text, first frames, and short camera-shot control.
Get Started
From a small-scale task to full integration.
01
Prepare the task and materials
Define objectives, required inputs, and output requirements, using real business examples as a starting point.
02
Try it in the API testing area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate according to the API documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Understand output quality and capability boundaries before formal use.
Ten-second generation cannot use all the controls available for five seconds: kling-v1's first and last frames and camera movement apply only to five-second tasks. The last frame can only be used for image-to-video and must be provided together with the first frame; if the shot must end on a specified image, plan it first as a five-second workflow.
kling-v1 does not support native audio or 4K mode, nor does it support Omni multi-image references or reference video editing. Talking photos require existing audio and are created through photo animation and lip synchronization; they cannot replace the generation of music, dialogue, or ambient sound.
Extension tasks do not support negative_prompt or cfg_scale, so not all control parameters from the initial generation can be reused unchanged. For photo lip-sync tasks, use a clear front-facing image of one person; audio duration should not exceed the selected video duration to avoid mismatches between source material and delivery length.
Frequently Asked Questions
Answers to common questions about using kling-v1.
How should I choose between five-second and ten-second tasks in kling-v1?
Choose five seconds when you need first and last frames or camera movement control; this is a clear functional boundary of this model. Ten seconds is suitable for generation tasks that do not rely on these two controls. Duration should be determined by the shot content; if you need to continue an existing clip, you can use video extension rather than assuming arbitrary lengths are directly supported.
Can I provide only an end frame for image-to-video?
No. Image-to-video requires start_image_url. The end frame end_image_url is only an optional ending reference and must be used together with the first frame. When using first and last frames with kling-v1, you must also choose a five-second task; with only a first frame, you can describe the actions you want to appear around this image.
How can I make generated content better match the prompt?
For initial generation, you can adjust prompt relevance through cfg_scale, ranging from 0 to 1; higher values adhere more closely to the prompt. negative_prompt can specify content you do not want to appear, up to 200 characters. Neither applies to extend; when continuing, focus on describing the actions in the next segment.
How can I make a photo speak?
Use /kling/talking-photo, submit image_url and audio_url, and explicitly set model to kling-v1. Supported audio formats are mp3, wav, m4a, and aac, with a maximum size of 5MB; videos can be five or ten seconds. It synchronizes lip movements with existing audio and does not create dialogue on its own.
How can generated results be passed to an application for further processing?
Video tasks return information such as video_url, video_id, task_id, and status: the video link is used to retrieve the finished video, the video ID is used for subsequent extension, and the task ID is used to track processing progress. Setting async=true lets you obtain the task ID first and then query the result; you can also configure a callback address to receive completion information.
Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.