How should I choose between kling-v1's five-second and ten-second tasks?
Choose five seconds when first/last frame or camera movement control is needed; this is a clear functional boundary of this model. Ten seconds is suitable for generation tasks that do not depend on these two controls. The duration should be determined by the shot content; if you need to continue an existing clip, use video extension rather than assuming arbitrary lengths are supported.
Can I provide only the last frame for image-to-video?
No. Image-to-video requires start_image_url; the last frame, end_image_url, is only an optional ending reference and must be used together with the first frame. When using first and last frames with kling-v1, you must also choose a five-second task; when only using the first frame, describe the desired action around that image.
How can I make generated content more closely match the prompt?
For initial generation, cfg_scale can adjust prompt relevance, ranging from 0 to 1; higher values adhere more closely to the prompt. negative_prompt can specify content you do not want to appear, up to 200 characters. Neither applies to extend; when continuing, focus on describing the action in the next segment.
How can I make a photo speak?
Use /kling/talking-photo, submit image_url and audio_url, and explicitly set model to kling-v1. Supported audio formats are mp3, wav, m4a, and aac, with a maximum size of 5MB; videos can be five or ten seconds. It synchronizes lip movements using existing audio and does not create dialogue on its own.
How can generation results be passed to the application for further processing?
Video tasks return information such as video_url, video_id, task_id, and status: the video link is used to retrieve the finished video, the video ID is used for subsequent extension, and the task ID is used to track processing progress. Setting async=true lets you obtain the task ID first and then query the result; you can also configure a callback address to receive completion information.