Fish Audio S1 is a text-to-speech model suited to courses, chapter narration, and continuously updated explanatory content. The platform guide positions it as a stability-focused choice, especially for production workflows that first make long text read smoothly, then adjust timbre and pacing. After specifying model: s1, you can reuse an existing voice or synthesize speech with a reference recording, and deliver results through an audio link.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and interface features
Clarify the model, inputs and outputs, and invocation method before choosing.
mp3, wav, pcm; both wav and pcm return a WAV container
Voice reference
reference_id or a single references sample; the two are mutually exclusive
Prosody control
prosody.speed controls speaking speed, and prosody.volume controls volume gain
S1 uses (parenthesis) emotion expressions; configure them separately from the [bracket] syntax of the S2 series. Capability descriptions combine public model materials and this platform's documentation; refer to the relevant API and pricing sections for actual parameters, outputs, and billing rules.
Core capabilities
Learn what voice tasks s1 is suited to solve.
Start with stable narration
Courses and long-form voiceovers need to maintain sentence relationships and vocal consistency throughout. The platform guide recommends comparing S1 for tasks that prioritize long-text stability; during production, you should still split content by chapter and review proper nouns and numbers, rather than treating model positioning as a guarantee that every output will be error-free.
Use a fixed voice for continuously updated content
Reuse a saved voice through reference_id so that openings, main content, and supplemental material in the same series retain a unified vocal identity. Voice references and model selection are configured separately, making this suitable for content projects with a fixed announcer or brand voice.
Bring speech into the editing workflow
You can choose MP3 delivery or use a WAV container for mixing and editing workflows. prosody controls speaking speed and volume, making it easy to generate a short sample first, establish the pacing, and then arrange subsequent chapters with the same settings.
Applicable Scenarios
Choose based on specific content and delivery methods.
Courses and Knowledge Explanations
Split teaching scripts into sections by topic, and preprocess abbreviations, formulas, and technical names. Use the same voice and speech rate to generate each section, then complete course production with subtitles and visuals after listening checks.
Audio Chapters and Long-Form Narration
Store scripts, models, and voice information by chapter. When pronunciation or pause issues are found, redo only the sections that need revision, reducing rework for the entire piece while maintaining the editing pace.
Continuously Maintained Application Prompts
Convert product tutorials, operating instructions, and help copy into a consistent voice. Keep version records for newly added or modified text, making it easier to reproduce the corresponding prompts after product updates.
How to Choose This Model
Compare based on scripts, voices, and production costs.
Prioritize Stable Reading
If the task focuses more on chapter continuity and repeatable editing, start by auditioning S1. Short content with stronger expressiveness can also be compared with S2 Pro; upgraded models do not replace evaluation using actual scripts.
Make a Clear Choice, Keep Versions
The platform's default model is s2-pro; using S1 requires setting the model request header. Saving the model, voice, format, and script version helps reproduce the production settings for a batch of content.
Getting Started: Keep Long Chapter Readings Clear and Consistent
Arrange the input first, then connect it to the corresponding application workflow.
Prepare Input
Prepare scripts by chapter, retaining paragraphs, punctuation, numbers, and uncommon terms; choose a fixed voice authorized for use.
Organize Calls and Subsequent Workflows
Submit text to /fish/tts, select s1 through the model request header, and provide either reference_id or a single recording reference. After completion, listen using audio_url; for series content, save the model, voice, speech rate, and format.
Practical Task Example: Keep Long Chapter Readings Clear and Consistent
Design tasks directly from the following inputs and acceptance priorities.
Recommended Task
Specify s1 in the model request header, generate a short sample from a paragraph mixing long and short sentences, and after confirmation produce it in sections using the same settings.
Key Checks
Check the beginnings and ends of sections, proper nouns, and voice consistency across sections; S1's parenthetical emotion syntax cannot be used directly as S2's square-bracket controls.
Usage Limits
Learn about synthesis methods and delivery scope.
This entry is for text-to-speech and is not equivalent to speech recognition, music generation, or native real-time audio streaming. Long-form content should be produced in segments and reviewed, checking numbers, abbreviations, and proper nouns.
One-time cloning accepts only one publicly accessible HTTPS MP3/WAV recording sample and its accurate verbatim transcript. It does not accept Base64, data URIs, or URLs with credentials; long-term voice profiles are not automatically saved.
format=pcm returns a WAV container and cannot be processed directly as raw PCM bytes; opus is not supported. For the supported range of generation parameters and actual billing, see this platform's API documentation and pricing.
Frequently Asked Questions
Answers to common questions about using s1.
Why choose S1?
The platform guide recommends S1 as an option focused on long-text stability, suitable for content that prioritizes clear reading. Actual performance depends on the script, voice reference, and generation settings, so first preview representative passages.
Where should the model be specified?
Specify s1 through the HTTP request header model. If not specified, the default is s2-pro; the voice reference_id belongs to voice selection and is configured separately from the synthesis model.
How should I choose between saving a voice and one-time cloning?
For fixed characters or series content, use reference_id to reuse a voice; for one-off projects, use references to provide a recording and transcript. The two are mutually exclusive, and one-time cloning accepts only one reference sample.
How can I control speech speed and output format?
prosody.speed=1.0 indicates the original speed, and volume uses dB, where 0 means no volume change. You can choose mp3, wav, or pcm; the latter two both use WAV containers, and MP3 bitrates can be 64, 128, or 192.
How can I track completion results for long scripts?
After setting callback_url, first save task_id and started_at, then wait for the completion callback; you can also query by task ID. Only after obtaining audio_url can you proceed to playback or editing.