- Model selection
- HTTP request header model: s2-pro
- Platform API
- POST /fish/tts; text is a required input
- Audio delivery
- Returns audio_url synchronously; supports callback_url asynchronous callbacks
- Output formats
- mp3, wav, pcm; both wav and pcm return a WAV container
- Voice reference
- reference_id or a single references sample; the two are mutually exclusive
- Prosody control
- prosody.speed controls speech rate, and prosody.volume controls volume gain
The S2 series can use [bracket] natural-language expression prompts in text, such as [whisper]; actual results should be confirmed by listening. Capability descriptions are based on public model materials and this platform's documentation; actual parameters, outputs, and billing rules are subject to the corresponding API and pricing sections.