Fish Audio S2.1 Pro
Fish Audio S2.1 Pro text to speech — stateless voice cloning, billed on input length.
Production multilingual speech with stateless voice cloning. Attach a short reference sample to the request and the output mimics that voice — no enrolment step, nothing stored.
No credit card required. Takes 30 seconds.
2,400+
Developers
1.2M+
API calls served
100+
Endpoints
$0.01
Per call
Yep, that's it.
Playground coming soon — use the API endpoint directly with your API key.
Pricing
Per character: $0.0317/1K chars
Endpoint
/v1/media/queue
Strengths
Send a voice sample with the request and get that voice back. There is no voice to create, upload, or manage beforehand.
The sample travels with the request and is not persisted, which keeps the consent story simple.
Speaking style is steered with open-ended instructions rather than a fixed voice list.
Built for expressive narration and multi-speaker dialogue across languages.
Quick start
Copy this snippet and start making calls with Fish Audio S2.1 Pro.
// Step 1: Submit job
const res = await fetch('https://api.yepapi.com/v1/media/queue', {
method: 'POST',
headers: {
'x-api-key': 'YOUR_API_KEY',
'Content-Type': 'application/json',
},
body: JSON.stringify({
"model": "fish-audio/s2.1-pro",
"prompt": "This line is spoken in a cloned voice, generated from a single reference sample sent with the request."
}),
});
const { data } = await res.json();
const jobId = data.jobId;
// Step 2: Poll for result
const status = await fetch(`https://api.yepapi.com/v1/media/status/${jobId}`, {
headers: { 'x-api-key': 'YOUR_API_KEY' },
});
const { data: job } = await status.json();
// job.status: "pending" | "processing" | "completed" | "failed"
// job.result: { text?, image?, audio?, video? }Why use Fish Audio S2.1 Pro through YepAPI?
Fish Audio S2.1 Pro API — voice cloning without enrolment
S2.1 Pro is Fish Audio's production text-to-speech model and the only speech model on YepAPI that supports stateless voice cloning. Attach a short reference sample to the request and the generated speech mimics that voice.
Because cloning is stateless there is no voice to create, upload, or manage — the sample travels with the request and nothing is retained. Billing is $0.0317 per 1,000 characters, cloned or not.
Start generating in 30 seconds
$5 free credit on signup. No credit card required. Pay per call.
What developers say
“Switched from SerpAPI and cut our SERP costs by 80%. Same data quality, way simpler billing.”
“One API key for AI models, SERP data, and web scraping. Saved us from managing 4 separate providers.”
“The $5 free credit let us prototype our entire rank tracking feature before committing. No other API does that.”
Frequently asked questions
Production multilingual speech with stateless voice cloning. Attach a short reference sample to the request and the output mimics that voice — no enrolment step, nothing stored.
Pricing for Fish Audio S2.1 Pro through YepAPI is based on usage. No monthly minimums — you only pay for what you use.
Sign up for a free API key, then send requests to the /v1/media/queue endpoint.
Ready to use Fish Audio S2.1 Pro?
$5 free credit on signup. No credit card required. Pay per call.
Explore more models
Fish Audio S2 Pro
Fish AudioFish Audio S2 Pro text to speech — expressive multi-speaker narration, billed on input length.
Fish Audio S1
Fish AudioFish Audio S1 text to speech — inline emotional control, billed on input length.
Fish Audio S2.1 Pro Free
Fish AudioFish Audio S2.1 Pro Free text to speech — free tier for prototyping, billed on input length.