Is digitalhuman a general-purpose text-to-video model?
It is mainly used for lip-synced talking based on face images or videos. The text field is for narration content and should not be understood as being able to generate arbitrary videos solely from scene descriptions. To create complex environments, shots, or storylines, choose a more suitable creation method.
Must the character material be a video?
It does not have to be limited to video; digital human creation also supports starting from face images. The interface provides image_url and video_url respectively, which can be prepared according to the material type. For first-time use, it is recommended to make a short test clip first, confirm the character's performance, and then use the same material for subsequent tasks.
Can I use my own recording or script?
The interface provides audio_url, as well as text and voice_id fields, allowing creation based on recordings or scripts. The two approaches should not be mixed without validation; first choose the narration method, test the corresponding material combination, and then use the successful configuration for ongoing production.
Can I clone a voice directly in the generation request?
The accompanying MCP tool supports cloning a voice using a short reference audio clip, but the video generation interface does not have a dedicated field for cloning samples. Prepare the voice separately first, then arrange talking-head generation; do not confuse character-driving audio with voice-cloning reference material.
How do I obtain the video after generation?
Submit a task through POST /digital-human/videos; the response provides the video URL and task-related fields. When using an asynchronous workflow, you can use callbacks or MCP task queries to obtain progress, and determine from the returned status whether a usable result is available before downloading the video for post-production.