Why choose S1?
The platform guide recommends S1 as an option focused on long-text stability, suitable for content that prioritizes clear reading. Actual performance depends on the script, voice reference, and generation settings, so first preview representative passages.
Where should the model be specified?
Specify s1 through the HTTP request header model. If not specified, the default is s2-pro; the voice reference_id belongs to voice selection and is configured separately from the synthesis model.
How should I choose between saving a voice and one-time cloning?
For fixed characters or series content, use reference_id to reuse a voice; for one-off projects, use references to provide a recording and transcript. The two are mutually exclusive, and one-time cloning accepts only one reference sample.
How can I control speech speed and output format?
prosody.speed=1.0 indicates the original speed, and volume uses dB, where 0 means no volume change. You can choose mp3, wav, or pcm; the latter two both use WAV containers, and MP3 bitrates can be 64, 128, or 192.
How can I track completion results for long scripts?
After setting callback_url, first save task_id and started_at, then wait for the completion callback; you can also query by task ID. Only after obtaining audio_url can you proceed to playback or editing.